US ATLAS Tier 2 Technical

US/Eastern
Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), Shawn Mc Kee (University of Michigan (US))
Description

Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.

Zoom Meeting ID
67453565657
Host
Fred Luehring
Useful links
Join via phone
Zoom URL
    • 11:00 11:10
      Introduction 10m
      Speakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))

      News:

      • Keep the capacity and services spreadsheets updated. Keep CRIC and OSG topology updated when servers are added or retired.
      • Ongoing review of FY26 and end-of-CA funding
        • From John Hobbs: "I had a talk with Aamir yesterday. He said that it's "highly, highly probable" that we'll get the full year 5 amount."

      Upcoming meetings:

      Unchanged from last meeting

      Open tickets:

      • Infrastructure tickets?
      • ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
      • ggus:3559 SWT2/OU: Dual-stack [on hold]
        • From Horst today: "The only one missing now is the CE, grid1.oscer.ou.edu"
      • ggus:1001382 TW-FTT: failing transfers as SOURCE due to certificate issue

      Operations:

      • AGLT2

      • MWT2

      • NET2

      • SWT2/CPB

      • SWT2/OU

       

    • 11:10 11:20
      TW-FTT 10m
      Speakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
      • Generally, site is running smoothly.
      • ggus:1001382 : There are no new updates.
    • 11:20 11:30
      AGLT2 10m
      Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)

      Fred noticed mystery hepscore value in facility spreadsheet of 17.48 for R6625 with AMD 9354
        Not clear where that value comes from
        UM added 12 R6625s in 2024Q3 and MSU added 5 nodes 2025Q3
        Measured hepscore is 22.57 (4 runs each on 3 msu and 2 um WNs)
        Updating spreadsheet changed AGLT2 average 13.91 -> 14.54 i.e. +4.53%
        Updated capacity tabs for 26Q1, 25Q4, 25Q3 but not able to change older tabs
        Updated atlas CRIC and OSG topology
        todo: submit update for hepix table 

      Also testing effect of vulnerability mitigation and updated CPU microcode 
        AMD 9354: 2.1% degradation in hepscore from mitigations and new microcode 
        Now testing some older CPU    
            
      Finished reconfiguring storage raid6 pools
        which had default 64k stripe size instead of desired 512k for better IO performance 

      Status of Fireflies/SciTags in dCache
         11.2.3-1 has almost all patches submitted by Shawn; functional but...
          TPC transfers are missing the start marker which would specify the activity type
          Transfers are still recorded via end markers but use default activity
          non-TPC transfers don't specify activity and use default anyway

      Added Kafka support in our dcache 
          with a kafka cluster on our 3 zoo nodes
          early problems with flooding of the log partitions,
          fixed it by reconfiguring it to limit the log size to 20GB

    • 11:30 11:40
      MWT2 10m
      Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • Scheduled dCache upgrade for 04/27
      • Working on procurement
      • Was set offline today (04/01) by HammerCloud. Still investigating, but most likely a pilot problem
      • GGUS ticket (1002243) regarding the RSE basepath prefixes. Will respond
    • 11:40 11:50
      NET2 10m
      Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

      Tape transfer errors, ggus 1002113,

      Tape Transfer Issues Report

      Two separate issues contributed to the tape transfer failures observed over the past few weeks.

      First, the space reservation for the tape buffer was being used in the SRR space usage report instead of the corresponding pool group. As a result, the space occupied by recalled files was not being accounted for correctly. This led Rucio to continue requesting additional files for transfer to tape under the false assumption that sufficient free space was still available. This issue has now been corrected.

      Second, following the tape backend firmware upgrade, new problems appeared and some recalls became very slow. This created complications for the new bringonline+transfers model that NET2 started using at the request of DDM Ops. In cases where a recall completed only after the associated FTS transfer had already timed out or been cancelled, the cancellation was not propagated back to dCache. The recalled file therefore remained pinned for several days as an orphaned file. Since no transfer was available to move those files out of the tape buffer, the buffer filled rapidly. This issue is still under investigation, but we are now monitoring these events closely.

      Unscheduled cluster downtime because the host certificate expired, and then because the certificate harvester uses to communicate with the cluster had expired.  This was followed by a brief blacklisting owing to a black hole node, which was removed from the cluster (it had developed a network issue during the downtime).

      Short scheduled cluster downtime for OS upgrade on the main router.

    • 11:50 12:00
      SWT2 10m
      Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))

      SWT2_CPB:

      • We rebuilt five R740 storage servers from EL7 to EL9 while preserving data. 

        • After the rebuilds, we verified the data and it appears to have been preserved. We have a temporary backup of the data in case of any data loss. 

        • The servers have been returned to production, and we have not observed any issues so far.

        • We are currently creating backups of additional R740 storage servers for migrating them from EL7 to EL9. 

      • We changed the certificate on one of our XRootD Proxy servers to make it a full-chain certificate. All of our XRootD Proxy servers now have full-chain certificates. 

        • We are monitoring for any new or increased errors such as with SSL Handshake Exception.

      • 11.9% of our production jobs have failed within 12 hours on 3/30 (457 failed production jobs).

        • One worker node (compute-19-39) showed 1% of jobs succeeding (6 succeeded, 378 failed). 

        • Removed this node from service for now to investigate. Ensuring there is not an issue with the worker node before returning putting it back into production. 

        • Specific Errors:

          • Failed to execute payload:PyJobTransforms.transform.execute CRITICAL Transform executor raised TransformValidationException: EVNTtoHITS got a SIGSEGV signal (exit code 139)

          • Non-zero exit code from transform substep executor

        • The errors started to increase at 2:20 a.m. UTC 3/30. So far, the peak was at 6:00 a.m. UTC, and lowered back to down at 8:40 a.m. UTC 3/30.

      OU:

      • Site running well
      • Network monitoring: waiting for OneNet network folks to respond
      • xrootd migration: need to re-mount OURdisk partition in a different way to ensure usage monitoring, then will start migration
      • Dual stack: only grid1 missing; also, waiting for additional OFFN ipv4 addresses
      • New SLURM version with cgroups v2 support: will schedule maintenance for that soon