US ATLAS Tier 2 Technical

US/Eastern
Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), Shawn Mc Kee (University of Michigan (US))
Description

Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.

Zoom Meeting ID
67453565657
Host
Fred Luehring
Useful links
Join via phone
Zoom URL
    • 11:00 11:10
      Introduction 10m
      Speakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))

      Happy new year!

      News:

      • Thanks to those who submitted the Operation Plan in December: they can be found here.
      • Capability mini-challenge planned for January 2026 (planning document here)
        • Host tuning: Planning document here. Second meeting this Friday (1/23), at 11am Eastern. Current T2 participants AGLT2, MWT2, and NET2.
        • Rucio/SENSE: current T2 participants MWT2 and NET2.
        • IPv4 blackout: scheduling work meeting (see newdle here) for the next few days. Current T2 participants: AGLT2 and NET2.
      • Keep the capacity and services spreadsheets updated.
      • Keep CRIC and OSG topology updated when servers are added or retired.
      • Quarterly reports are due ASAP!
      • We are trying to schedule our next procurement meeting for Friday (2/6), at 11am Eastern. More details to come.
      • NSF CA review next week. Possible homework on Wednesday evening.

      Upcoming meetings:

       

      • Varnish meeting [internal ATLAS meeting, but interesting for this group: Jan 26th at CERN]

      Discussion during the meeting:

      • New dCache version released today Release 11.2.X [release notes here].
      • Code for testing cgroup memory killing here (from Aidan)
      • AGLT2 notes that default dCache strip size (64k) is suboptimal. Both AGLT2 and MWT2 are using 512k for better IO performance. All sites should check.

      Open tickets:

      Operations:

      • AGLT2

      • MWT2

      • NET2

      • SWT2/CPB

      • SWT2/OU

    • 11:10 11:20
      TW-FTT 10m
      Speakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))

       GGUS 1001382 : Due to the asymmetric route, we are negotiating with the network provider.

    • 11:20 11:30
      AGLT2 10m
      Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)

      We had noticed some pool nodes with lower IO performance during the last mini data challenge.
        Now sending sar monitoring to our logstash and opensearch to help identify the slower file systems.
        The suspected cause is a smaller strip size (64K or 128K),
           smaller than the 512K we had determined to be our chosen compromise.
        We will need to drain and recreate the RAID arrays on these nodes.

      On 1/16, We noticed some pilot:1361 errors for input file access failures for 44 jobs.
        All errors came from one dCache pool node
        In hindsight, it seemed to have recovered by itself, before we restarted the dCache services on that node.
        This was one of the nodes already identified with lower IO performance.

    • 11:30 11:40
      MWT2 10m
      Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • UIUC had a PM on 01/07
      • Finished our quarterly report
      • An IU switch to be rebooted on 01/26
      • Plan to upgrade dCache when the next golden release comes out
      • Updated accounting in topology

       

    • 11:40 11:50
      NET2 10m
      Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

      GGUS 3255: closed, we see only occasional backoff limit failures so we're confident that the underlying issue with the pilot not seeing the right amount of disk has been fixed.

      GGUS 1001575: due to an issue with the tape system throughput.  The staging errors have gone away but the throughput issue is still there, we are following up with NESE.

      GGUS 1001545: probably due to an issue with updating packages on our storage servers, NESE was migrating to a new proxy for this purpose.  The migration is now finished so the updates can go ahead.

      A couple of minor blacklisting episodes, one related to a surge of merge jobs and one related to the proxy issue above.  Also a site draining episode apparently due to a connection issue with UVic.

    • 11:50 12:00
      SWT2 10m
      Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
      • After finishing our troubleshooting with EL9 XRootD Proxy performance issues, we updated the rest of our XRootD proxy servers to XRootD 5.7.2 to finish fixing these issues. All four now appear to be working as expected.

      • We met with Facilities Management on 12/10 and performed a partial, intential drain on 12/31 in preparation for the power switchover on 1/2. 

        • Unfortunately, we lost power on 1/2, 1/3, and 1/5 for different reasons during the switchover to a new rental generator, which was needed for UTA Facilities Management to perform future power work

        • On 1/2, the UPS battery runtime estimate was very inaccurate and misled us about how much time we had for Facilities Management to complete the switchover. 

        • On 1/3, the switchover to the rental generator was completed, but the UPS continued to drain. We powered off our storage right before losing power in an attempt to protect data. When the UPS attempted to power back on, it generated smoke. We powered it down, then placed it into bypass mode. 

        • On 1/5, a Schneider engineer investigated the UPS issue. During that work, the rental generator powered off and we were restored to building power. We confirmed the rental generator was not the correct type for our power needs and that the UPS was not compatible with it.

      • After we recovered from the 1/5 outage and brought systems back online, we experienced transfer efficiency issues related to DNS. 

        • First, our internal DNS service was in an abnormal state after the restart and was not functioning properly (e.g., nodes could not reach external repositories to install specific packages for testing). 

        • Second, we needed to update DNS to include our newly rebuilt EL9 DTN servers. 

        • After making the DNS updates and restarting the service, transfer efficiency returned to normal.

      • Campus Facilities Management successfully tested the building generator on 1/7, confirming it will automatically switch over in the event of another power outage.

      • We improved our process for rebuilding storage without migrating data to another system. We tested this and rebuilt a production storage server from EL7 to EL9. No data was lost. As a precaution, we also copied the data to an old storage server that will serve as a temporary storage area for future rebuilds.

      • Because one of our Frontier Squid servers (slate01) was not functioning properly and is not necessarily required, we disabled it and removed it from our site’s proxy list in CRIC.


      OU:

      • Running smoothly
      • Working on new xrootd servers, need Hiro's help
      • Have new test SLURM jobmanager up, need to test cgroups v2 memory killing