US ATLAS Tier 2 Technical

US/Eastern
Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), Shawn Mc Kee (University of Michigan (US))
Description

Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.

Zoom Meeting ID
67453565657
Host
Fred Luehring
Useful links
Join via phone
Zoom URL
    • 11:00 11:10
      Introduction 10m
      Speakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))

      Quick links:

      General news

      • News about equipment purchase are expected this week.
      • Still facing serious supply chain problems (several vendors refusing purchase orders)
      • In discussions with management about the situation.

      Upcoming meetings:

      Open tickets:

    • 11:10 11:20
      TW-FTT 10m
      Speakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
      • We had IPv6 connectivity problems between Aug. 11 and Aug. 13.
    • 11:20 11:30
      AGLT2 10m
      Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)

      No significant incident over last 2 weeks

      Met with Dell sales and server representatives on Aug 18.
      Dell said 16G DIMMs are not available.
      So we will need to order systems with 32G DIMMs.
      We will investigate the impact on HS23 after populating
      only half the slots with HT ON and quarter slots with HT OFF
      using our most recent AMD worker nodes (R6625 AMD 9354).

    • 11:30 11:40
      MWT2 10m
      Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • GGUS:1003511: A/R recalculated for June and July
        • All six gatekeepers were drained for short periods of time for CVE kernel updates
        • Two of our six gatekeepers were also drained for an extended period of time for kernel updates
        • The site was online and full throughout the maintenance, but the drained gatekeepers caused the reporting to fail
      • Working with both IU and UC networking on network quotes
    • 11:40 11:50
      NET2 10m
      Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

      We are working with Dell to figure out what the problem is with the dense nodes, which see many jobs fail due to backoff errors.  We removed them from the cluster for some firmware upgrades, which helped, but we still see backoff errors, so we are continuing to discuss this with them.

      The tape has been in downtime due to issues with the throughput.  A software upgrade to the tape system was carried out on Monday, we are checking to see if it has solved the problem.

    • 11:50 12:00
      SWT2 10m
      Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))

      SWT2_CPB: 

      • One of our CRAC units is experiencing issues. It is due to an issue with the fan motor. 

        • Facilities Management powered this unit off on 8/6. 

        • The replacement part is expected to arrive on 8/19. 

        • To reduce temperature, we drained our oldest worker nodes (R410s) on 8/7. Temperature in the room appears to be low enough and stable after this. We will put these nodes back into service once repairs are complete. 

      • Two of our R740 storage servers had failed OS drives. One also had a failed fan. We worked with the vendor and had them repaired. 

        • We migrated data from these nodes to other storage, so that powering them off for repairs would not cause disruptions. 

      • The SWT2_CPB_K8S queue has been disabled in CRIC. 

        • We rebuilt the K8S worker nodes as worker nodes for the SWT2_CPB PQ. We will not put these into service until the problematic CRAC unit is fixed (1048 job slots). 

      • Concerning memory issues with our XRootD Proxy servers:

        • Adding jemalloc back to our XRootD Proxy servers has improved memory management significantly. 

        • The custom changes we made to the service unit has helped restart these services when they experience OOM events, resulting in quick recovery.

        • We are experiencing issues with one of the four XRootD Proxy servers and are investigating this. It has experienced  OOM events three times over the past five days, starting on 8/14. 

      • GGUS-Ticket-ID: #1003053 - IGTF CRLs & fetch-crl SHA1 signature validation (SWT2_CPB)

        • We have not updated the ticket, but we are closely monitoring the CRL update on our servers. 

      • GGUS-Ticket-ID: #1003573 (old ticket) and  1002282 (current)

        • Jobs continue to use shared/NAS dir for home directory instead of locally on worker nodes. 

        • We attempted to change home dir on worker nodes, but this did not work. 

        • It may need to be fixed at pilot level. 


      OU:

      • OSCER Maintenance today
      • Will update se1's osg-ca-certs as well
      • Storage migration: all rucio tests are succeeding, but still need help migrating the SRR python script