US ATLAS Tier 2 Technical

US/Eastern
Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), Shawn Mc Kee (University of Michigan (US))
Description

Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.

Zoom Meeting ID
67453565657
Host
Fred Luehring
Useful links
Join via phone
Zoom URL
    • 11:00 11:10
      Introduction 10m
      Speakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))

      Quick links:

      General news

      • Ongoing discussion about equipment funding based on new end-of-CA estimate ongoing.
      • Still facing serious supply chain problems (several vendors refusing purchase orders)
      • In discussions with management about the situation.

      Upcoming meetings:

      Open tickets:

      • ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
      • ggus:1003053 SWT2/CPB: IGTF CRLs & fetch-crl SHA1 signature validation
    • 11:10 11:20
      TW-FTT 10m
      Speakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
      • All jobs failed for approximately 6 hours on Aug. 30th, should be caused by network switch failure.
    • 11:20 11:30
      AGLT2 10m
      Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)

      Starting to test RHEL10
        built a test pool node VM to work on ansible updates
        Plan to have some fraction of the newer dCache pool nodes on RHEL10
          for comparison during next capacity challenge
        built a test WN VM, it finished 18 ATLAS jobs and 3 failed. 

      S3 storage problem
          Certificate expired on Aug 7th,
          Replaced it with an IGTF cert we had pre-emptively renewed before InCommon transitioning to CERTinext,
          Still working on access to new IGTF/OV cert from InCommon/CERTInext    

    • 11:30 11:40
      MWT2 10m
      Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • UIUC had a short network maintenance on August 27th
      • Some network instability at UC on August 29th put us offline multiple times. It was resolved by Sunday
      • Was put offline on August 31st most likely from the Analysis Facility downtime. It was resolved by afternoon. Will investigate further
    • 11:40 11:50
      NET2 10m
      Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

      Blacklisted yesterday with everybody else.  There was also a second dip in the number of running slots around 4:30 am Eastern this morning, shared by all US sites, do we know what caused that?

      Making progress on limiting backoff errors, traced the nvme errors we were seeing to an issue with the kernel, which we have a workaround for.  Additionally, the newer OKD version enabled more compression of resources under load by default, meaning that the nodes could run at 15-20% over CPU capacity.  This turned out to be too much, so we turned that off.  Over the last couple of days the number of backoff errors seen is much reduced.  We are still running at a lower amount of slots because of issues with the two machines that required disk replacements.

      The tape is out of downtime, we are testing to see if the throughput has improved with the fixes that were made.

    • 11:50 12:00
      SWT2 10m
      Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))

      SWT2_CPB: 

      • Investigated reported issues concerning deletions by a user. 

        • This may be an issue involving DDM, but discussions are ongoing. 

        • We are working toward understanding this issue better. 

      • jemalloc has significantly helped with memory management on our XRootD Proxy servers. 

        • Two of our XRootD Proxy servers still restart their service from OOM events from time to time, but less frequently than before using jemalloc. 

        • The changes we made to the service to help protect against significant disruptions is working as intended. Restarts are quick, minimizing transfer disruptions during OOM events and failed service. 

        • Removed one of our XRootD Proxy servers temporarily to investigate hardware issues and other possible causes of memory issues. 

          • We have added this server back into service. 

      • Testing new configuration for perfSONAR machines with current plans to rebuild our perfSONAR machines with this new configuration for improved security. 

        • New planned build is testpoint with containers.

        • Thanks for the quick patch from perfSONAR developers to the security issue identified at SWT2.
        • The developers had additional suggestions for all US ATLAS sites - we should follow up in a dedicated meeting.
      • Campus networking has performed a network upgrade in the past two weeks. 

        • The upgrade doubled the bandwidth of the bundle connecting the machine room from 40 to 80 Gbps. This is part of a multi-step phased networking upgrade over the next year.

        • We are still monitoring to confirm this change. 

        • They attempted the upgrade on 8/21. There were issues with transfers. They responded quickly by disabling the new connection to investigate. 

        • They found the issue, resolved it, then added the new connection on 8/29. We monitored and there were no issues. 

      • Our problematic CRAC unit was repaired and powered back on. We monitored then put our drained worker nodes back into service on 8/24 and 8/25. 

      • Previous issues: 

        • GGUS-Ticket-ID: #1003053 - IGTF CRLs & fetch-crl SHA1 signature validation

          • Problem with connect between one of our DTNs and BEIJING site (check crl ) (we have not made any changes the past two weeks, still investigating, but we remember this issue)  

        • We have a few jobs running on the TEST-CE. We do not have hammercloud jobs on this cluster. 

        • From time to time we see the "Unknown" status of our CE. We do not understand reason now 


      OU:

      • Had a failed raid6 drive and some xfs file system corruption on one of the 7 old xrootd data servers. Fixed.
      • Migration to new storage is making progress. Getting up to 4 GB/s inbound WAN transfers to new storage now.
      • Waiting to hear back from Timo et al about switching Panda over to start using the new storage for stage-in/out.