US ATLAS Tier 2 Technical

US/Eastern
Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), Shawn Mc Kee (University of Michigan (US))
Description

Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.

Zoom Meeting ID
67453565657
Host
Fred Luehring
Useful links
Join via phone
Zoom URL
    • 11:00 11:10
      Introduction 10m
      Speakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))

      Quick links:

      General news

      • We need to come up with site-level milestones for the next FY. Inputs to Fred and Rafael are appreciated.
      • If sites have ideas about use of shared ESnet xcache, this will be discussed in the next S&C week,

      Upcoming meetings:

      Open tickets:

      Operations:

      • Site production during the previous 2 weeks: AGLT2MWT2NET2, SWT2 (CPBOU), TW
      • TW

        AGLT2

        MWT2

        NET2

        SWT2/CPB

        SWT2/OU

         

       

       

    • 11:10 11:20
      TW-FTT 10m
      Speakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
      • Generally, site is running smoothly.
    • 11:20 11:30
      AGLT2 10m
      Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
      • Corrected the missing accounting data for May
          Still trying to figure out the current double accounting issue
          
        - Testing RHEL9.8 on service nodes, one dcache pool node and one worker node.
          Waiting for kernel fix for ipv6 bug CVE-2026-43038 to update pool nodes
          and lustre support for 9.8 kernel to update UM worker nodes.
      • Chasing dcache billing errors: timeout on get operations
        • found 19 files which were not present in dcache but registered in RUCIO, we declared them as lost files
        • Had a misunderstanding about the get errors, initially thought they were READ accesses to the dCache and could not find these files registered in RUCIO to AGLT2, after more debug, found out they were actually failed writes. 
        • A couple of reasons led to the spike of failed writes: 1) transfers from a problematic site FREIBURG site 2)  a spike checksum mismatches after finishing transfers, sources varied. 
    • 11:30 11:40
      MWT2 10m
      Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • UIUC compute moving to new racks with power and cooling improvements. All back online now
      • Both UC and IU working with networking teams to build out and quote hardware for increased network capacity and hardware refreshes
        • UIUC network capacity is enough for the compute there
      • testing cvmfs-2.14.0~pre4-2 on a subset of older compute nodes. Not seeing any glaring issues after the update
      • Started discussing investigating migrating off deprecated systems that will need to be done before el10
        • such as iptables -> nftables, network-scripts -> NetworkManager, etc.
    • 11:40 11:50
      NET2 10m
      Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

      NET2 is facing a challenging issue with sites not using LHCONE.

      This plot shows our "commodity internet" WAN link, which is not supposed to receive data. This last week, we saw several peaks at ~20 Gb/s of data being transferred from UK and GE sites to NET2 using this link. This has a negative impact to other people using the same "commodity internet" link in our science DMZ and it's not acceptable. If this continue, we will have to rate-limit this connection to 10 Gb/s negatively impacting the NET2 transfer efficiency.

      We have been observing that the number of non-LHCONE transfers is increasing with time. These 20 Gb/s peaks used to be rare and this last week they became very common.

      Discussion:

      • Can we do something at the WLCG/ATLAS level to fix this?
      • Can we do something at the US ATLAS level to fix this?

       

      month

      Downtime from Sunday night through Thursday afternoon last week due to yearly maintenance period at the data center, started early to allow for some work to be done on the storage.

      GGUS 1002244 closed, basepath changed during the downtime.

      Coming out of the downtime, observed some issues in the dense servers, this is under investigation.  At the moment therefore we are running with reduced slots.

    • 11:50 12:00
      SWT2 10m
      Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))

      SWT2_CPB: 

      • We had campus networking create new mailing lists for us and added them to CRIC for the security and admin contacts fields. 

      • GGUS-Ticket-ID: #1003078: Transfer failures as source with File Not Found error

        • One of our R740xd2 storage servers experienced issues with its RAID controller.

        • We worked with the vendor per warranty to troubleshoot issues. 

        • We had an idea to drain a different R740xd2 server that had very little data, take the RAID controller from the server, then replace it in the problematic storage server and import configuration. The vendor agreed and allowed us to do this.

        • Once this was complete, we monitored it and it remained stable for days. We migrated data from this storage.

        • This storage also has a problematic drive, which we will replace soon. 

        • Two storage servers will be unavailable temporarily as we perform tests. 

        • The vendor is sending us a new RAID controller and new drive. 

        • The monitoring and alerts we have in place are working. 

      • Both of our CE (gk01 and gk10) were in the same rack. We drained gk01, moved it to another rack. 

        • Once gk01 was moved, we drained gk10 while allowing gk01 to start handling new jobs. 

        • We also performed these steps in preparation to upgrade a network switch in the same rack gk10 is located. 

        • We upgraded the switch for three of our four XRootD Proxy servers and will do the fourth one within two months. We will be testing and monitoring the current new switches to ensure the throughput increase is evident. 

        • We experience some transfer errors during these upgrades. 

      • GGUS-Ticket-ID: #1003053: IGTF CRLs & fetch-crl SHA1 signature validation

        • We have tried multiple troubleshooting steps for this and are continuing to work on it.

        • We performed steps in the ticket to modify crypto-policy for one of our four XRootD Proxy servers. The rest were fine. 

        • We tried switching from the full-chain to leaf-only certificate. 

        • We reviewed other information on this node. 

        • We are going to renew the certificate, monitor, then rebuild it if we continue to see issues.