US ATLAS Tier 2 Technical

US/Eastern
Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), Shawn Mc Kee (University of Michigan (US))
Description

Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.

Zoom Meeting ID
67453565657
Host
Fred Luehring
Useful links
Join via phone
Zoom URL
    • 11:00 11:10
      Introduction 10m
      Speakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))

      News:

      • Keep the capacity and services spreadsheets updated. Keep CRIC and OSG topology updated when servers are added or retired.
      • Genesis proposal due in another 2 weeks! Several projects related to the US ATLAS Tier 2 facilities.
      • Quarterly reports due on Friday.

      Upcoming meetings:

      Unchanged from last meeting

      Open tickets:

      Unchanged from last week

      Operations:

      • AGLT2

      • MWT2

      • NET2

      • SWT2/CPB

      • SWT2/OU

       

    • 11:10 11:20
      TW-FTT 10m
      Speakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
      • Several worker nodes were added, increasing the total to 4272 slots.
      • A decrease in the number of running slots from April 8 to April 10 due to a lot of transferring. Removed fairshare policy
    • 11:20 11:30
      AGLT2 10m
      Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)

      Smooth running

      Waiting on dCache 12.2.4 which should include accepted last pull request
        12.2.3 was still missing sending packet at the start of the flow to identify flow/activity type

      Effect of vulnerability mitigations on hepscore 
        Had measured about 2% on AMD 9354; so not worth worrying about
        Do we need to ponder what would we do if it was 10,20,30% ?
        Quick scan of CPU generations purchased in last 10 years
        About 2% on 4x AMD CPUs and about 1% or less on 4x Intel CPUs

    • 11:30 11:40
      MWT2 10m
      Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • UIUC PM Today (April 15) - ~1/3 compute offline. Site otherwise running fine
      • Downtime scheduled for April 27 for dCache upgrade to 11.2.x and to update the RSE basepath for GGUS ticket #1002243
      • Fred to work with Dell on benchmark and testing various gen 5 CPU configurations
    • 11:40 11:50
      NET2 10m
      Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

      One short blacklisting due to problems with a storage pool, otherwise smooth running.

      Proposed fixing the basepath prefix without downtime but it seems that the experts prefer to wait for one.  Next foreseen downtime is in June, when MGHPCC does its yearly maintenance.

    • 11:50 12:00
      SWT2 10m
      Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))

      SWT2_CPB:

      • We rebuilt four storage servers from EL7 to EL9.

        • There were some transfer errors caused by these rebuilds. 

        • No data has been lost. Backups were made before rebuilds. 

        • We checked and verified data was not lost after rebuilds were complete. 

      • Our site experienced brief intermittent exclusion by HC.

        • This is a known issue to us that we believe we understand. 

        • We have multiple servers that are changing between read-only and read-write depending on the needs of our situation  with migrating and transitioning storage from EL7 to EL9. 

        • If the number of read-only storage servers are too high, this can cause other servers that are in read/write to have a high load. 

        • Once transitioning storage from EL7 to EL9 is complete, this issue should not occur again because less servers will be in this temporary read-only state. 

      • GGUS-Ticket-ID: #1002282 - Jobs Mistakenly Using Home Directory

        • We found that jobs are using the home directory on our NAS server, causing slower performance. We have a scratch directory on worker nodes dedicated for uses such as this.

          • Container files are being stored and referenced here during execution of jobs.

        • The directories created by jobs in this area on our NAS were not cleaned up, causing a buildup of these directories. 

        • After some discussion in the ticket, including some suggestions but also questions concerning our current working directory not being used for these container files, this area was cleaned up which did help reduce issues. The ticket was closed. However, we do not believe that this issue has been resolved. This requires having jobs stop using other directories except scratch. 


      OU:

      • Site running well
      • Had some storage overload because of large data influx and heavy I/O jobs; resolved itself
      • Network monitoring: still waiting for OneNet network folks to respond
      • xrootd migration: have created 1 PB OURdisk partition; need to re-mount in a different way to ensure usage monitoring, then will start migration
      • Dual stack: only grid1's ipv6 address missing; unfortunately, OSCER admin out sick currently
      • New SLURM version with cgroups v2 support: will schedule maintenance for that soon