US ATLAS Tier 2 Technical
Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.
-
-
11:00
→
11:10
Introduction 10mSpeakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
Quick links:
General news
- News about equipment purchase are expected this week.
- Still facing serious supply chain problems (several vendors refusing purchase orders)
- In discussions with management about the situation.
Upcoming meetings:
- 57th LHCOPN-LHCONE meeting [Oct 1st -- 2nd, Naples]
- Hepix Indico: HEPiX Fall 2026 Workshop (19-23 October 2026): Overview · Indico
- WLCG/HSF Workshop 2026 [Nov 2nd -- 6th, Bologna]
Open tickets:
- ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
- ggus:1003579 , ggus:1003580 , ggus:1003581 TW-FTT, NET2, and SWT2/OU need to update CA bundle.
- 11:10 → 11:20
-
11:20
→
11:30
AGLT2 10mSpeakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
No significant incident over last 2 weeks
Met with Dell sales and server representatives on Aug 18.
Dell said 16G DIMMs are not available.
So we will need to order systems with 32G DIMMs.
We will investigate the impact on HS23 after populating
only half the slots with HT ON and quarter slots with HT OFF
using our most recent AMD worker nodes (R6625 AMD 9354). -
11:30
→
11:40
MWT2 10mSpeakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
- GGUS:1003511: A/R recalculated for June and July
- All six gatekeepers were drained for short periods of time for CVE kernel updates
- Two of our six gatekeepers were also drained for an extended period of time for kernel updates
- The site was online and full throughout the maintenance, but the drained gatekeepers caused the reporting to fail
- Working with both IU and UC networking on network quotes
- GGUS:1003511: A/R recalculated for June and July
-
11:40
→
11:50
NET2 10mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
We are working with Dell to figure out what the problem is with the dense nodes, which see many jobs fail due to backoff errors. We removed them from the cluster for some firmware upgrades, which helped, but we still see backoff errors, so we are continuing to discuss this with them.
The tape has been in downtime due to issues with the throughput. A software upgrade to the tape system was carried out on Monday, we are checking to see if it has solved the problem.
-
11:50
→
12:00
SWT2 10mSpeakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
SWT2_CPB:
-
One of our CRAC units is experiencing issues. It is due to an issue with the fan motor.
-
Facilities Management powered this unit off on 8/6.
-
The replacement part is expected to arrive on 8/19.
-
To reduce temperature, we drained our oldest worker nodes (R410s) on 8/7. Temperature in the room appears to be low enough and stable after this. We will put these nodes back into service once repairs are complete.
-
Two of our R740 storage servers had failed OS drives. One also had a failed fan. We worked with the vendor and had them repaired.
-
We migrated data from these nodes to other storage, so that powering them off for repairs would not cause disruptions.
-
The SWT2_CPB_K8S queue has been disabled in CRIC.
-
We rebuilt the K8S worker nodes as worker nodes for the SWT2_CPB PQ. We will not put these into service until the problematic CRAC unit is fixed (1048 job slots).
-
Concerning memory issues with our XRootD Proxy servers:
-
Adding jemalloc back to our XRootD Proxy servers has improved memory management significantly.
-
The custom changes we made to the service unit has helped restart these services when they experience OOM events, resulting in quick recovery.
-
We are experiencing issues with one of the four XRootD Proxy servers and are investigating this. It has experienced OOM events three times over the past five days, starting on 8/14.
-
GGUS-Ticket-ID: #1003053 - IGTF CRLs & fetch-crl SHA1 signature validation (SWT2_CPB)
-
We have not updated the ticket, but we are closely monitoring the CRL update on our servers.
-
Jobs continue to use shared/NAS dir for home directory instead of locally on worker nodes.
-
We attempted to change home dir on worker nodes, but this did not work.
-
It may need to be fixed at pilot level.
OU:- OSCER Maintenance today
- Will update se1's osg-ca-certs as well
- Storage migration: all rucio tests are succeeding, but still need help migrating the SRR python script
-
-
11:00
→
11:10