US ATLAS Tier 2 Technical
Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.
-
-
1
IntroductionSpeakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
News:
- Keep the capacity and services spreadsheets updated. Keep CRIC and OSG topology updated when servers are added or retired.
- Genesis proposal due in another 2 weeks! Several projects related to the US ATLAS Tier 2 facilities.
- Quarterly reports due on Friday.
Upcoming meetings:
Unchanged from last meeting
- LHCOPN-LHCONE meeting #56 [Apr 15-16 in Montreal]
- Happening today and tomorrow!
- HEPiX Spring 2026 Workshop [Apr 20-24 in Lisbon]
- dCache workshop [May 6-7 at NIKHEF]
- CHEP 2026 [23-29 May in Bangkok, Thailand]
Open tickets:
Unchanged from last week
- Infrastructure tickets
- ggus:1002243 MWT2: RSE basepath prefix
- ggus:1002244 NET2: RSE basepath prefix
- ggus:1002233 AGLT2: RSE basepath prefix
- ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
- ggus:3559 SWT2/OU: Dual-stack [on hold]
- ggus:1001382 TW-FTT: failing transfers as SOURCE due to certificate issue
Operations:
- AGLT2
- MWT2
- NET2
- SWT2/CPB
- SWT2/OU
- 2
-
3
AGLT2Speakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
Smooth running
Waiting on dCache 12.2.4 which should include accepted last pull request
12.2.3 was still missing sending packet at the start of the flow to identify flow/activity type
Effect of vulnerability mitigations on hepscore
Had measured about 2% on AMD 9354; so not worth worrying about
Do we need to ponder what would we do if it was 10,20,30% ?
Quick scan of CPU generations purchased in last 10 years
About 2% on 4x AMD CPUs and about 1% or less on 4x Intel CPUs -
4
MWT2Speakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
-
5
NET2Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
-
6
SWT2Speakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
SWT2_CPB:
-
We rebuilt four storage servers from EL7 to EL9.
-
There were some transfer errors caused by these rebuilds.
-
No data has been lost. Backups were made before rebuilds.
-
We checked and verified data was not lost after rebuilds were complete.
-
Our site experienced brief intermittent exclusion by HC.
-
This is a known issue to us that we believe we understand.
-
We have multiple servers that are changing between read-only and read-write depending on the needs of our situation with migrating and transitioning storage from EL7 to EL9.
-
If the number of read-only storage servers are too high, this can cause other servers that are in read/write to have a high load.
-
Once transitioning storage from EL7 to EL9 is complete, this issue should not occur again because less servers will be in this temporary read-only state.
-
GGUS-Ticket-ID: #1002282 - Jobs Mistakenly Using Home Directory
-
We found that jobs are using the home directory on our NAS server, causing slower performance. We have a scratch directory on worker nodes dedicated for uses such as this.
-
Container files are being stored and referenced here during execution of jobs.
-
The directories created by jobs in this area on our NAS were not cleaned up, causing a buildup of these directories.
-
After some discussion in the ticket, including some suggestions but also questions concerning our current working directory not being used for these container files, this area was cleaned up which did help reduce issues. The ticket was closed. However, we do not believe that this issue has been resolved. This requires having jobs stop using other directories except scratch.
OU:- Site running well
- Had some storage overload because of large data influx and heavy I/O jobs; resolved itself
- Network monitoring: still waiting for OneNet network folks to respond
- xrootd migration: have created 1 PB OURdisk partition; need to re-mount in a different way to ensure usage monitoring, then will start migration
- Dual stack: only grid1's ipv6 address missing; unfortunately, OSCER admin out sick currently
- New SLURM version with cgroups v2 support: will schedule maintenance for that soon
-
-
1