US ATLAS Tier 2 Technical
Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.
-
-
11:00
→
11:10
Introduction 10mSpeakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
Happy new year!
News:
- Thanks to those who submitted the Operation Plan in December: they can be found here.
- Capability mini-challenge planned for January 2026 (planning document here)
- Host tuning: Planning document here. Second meeting this Friday (1/23), at 11am Eastern. Current T2 participants AGLT2, MWT2, and NET2.
- Rucio/SENSE: current T2 participants MWT2 and NET2.
- IPv4 blackout: scheduling work meeting (see newdle here) for the next few days. Current T2 participants: AGLT2 and NET2.
- Keep the capacity and services spreadsheets updated.
- Keep CRIC and OSG topology updated when servers are added or retired.
- Quarterly reports are due ASAP!
- We are trying to schedule our next procurement meeting for Friday (2/6), at 11am Eastern. More details to come.
- NSF CA review next week. Possible homework on Wednesday evening.
Upcoming meetings:
-
CHEP 2026 [23-29 May in Bangkok, Thailand]
- ATLAS S&C meeting [Feb 9-13 at CERN]
- WLCG Open Technical Forum (OTF) #8 on tape evolution [Feb 3-4 at CERN]
- LHCOPN-LHCONE meeting #56 [Apr 15-16 in Montreal]
- HEPiX Spring 2026 Workshop [Apr 20-24 in Lisbon]
- dCache workshop announced for May at NIKHEF. More details to come
- Varnish meeting [internal ATLAS meeting, but interesting for this group: Jan 26th at CERN]
Discussion during the meeting:
- New dCache version released today Release 11.2.X [release notes here].
- Code for testing cgroup memory killing here (from Aidan)
- AGLT2 notes that default dCache strip size (64k) is suboptimal. Both AGLT2 and MWT2 are using 512k for better IO performance. All sites should check.
Open tickets:
- ggus:1001575 NET2: staging error
- ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
- ggus:3559 SWT2/OU: Dual-stack [on hold]
Operations:
- AGLT2
- MWT2
- NET2
- SWT2/CPB
- SWT2/OU
-
11:10
→
11:20
TW-FTT 10mSpeakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
GGUS 1001382 : Due to the asymmetric route, we are negotiating with the network provider.
-
11:20
→
11:30
AGLT2 10mSpeakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
We had noticed some pool nodes with lower IO performance during the last mini data challenge.
Now sending sar monitoring to our logstash and opensearch to help identify the slower file systems.
The suspected cause is a smaller strip size (64K or 128K),
smaller than the 512K we had determined to be our chosen compromise.
We will need to drain and recreate the RAID arrays on these nodes.On 1/16, We noticed some pilot:1361 errors for input file access failures for 44 jobs.
All errors came from one dCache pool node
In hindsight, it seemed to have recovered by itself, before we restarted the dCache services on that node.
This was one of the nodes already identified with lower IO performance. -
11:30
→
11:40
MWT2 10mSpeakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
-
11:40
→
11:50
NET2 10mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
GGUS 3255: closed, we see only occasional backoff limit failures so we're confident that the underlying issue with the pilot not seeing the right amount of disk has been fixed.
GGUS 1001575: due to an issue with the tape system throughput. The staging errors have gone away but the throughput issue is still there, we are following up with NESE.
GGUS 1001545: probably due to an issue with updating packages on our storage servers, NESE was migrating to a new proxy for this purpose. The migration is now finished so the updates can go ahead.
A couple of minor blacklisting episodes, one related to a surge of merge jobs and one related to the proxy issue above. Also a site draining episode apparently due to a connection issue with UVic.
-
11:50
→
12:00
SWT2 10mSpeakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
-
After finishing our troubleshooting with EL9 XRootD Proxy performance issues, we updated the rest of our XRootD proxy servers to XRootD 5.7.2 to finish fixing these issues. All four now appear to be working as expected.
-
We met with Facilities Management on 12/10 and performed a partial, intential drain on 12/31 in preparation for the power switchover on 1/2.
-
Unfortunately, we lost power on 1/2, 1/3, and 1/5 for different reasons during the switchover to a new rental generator, which was needed for UTA Facilities Management to perform future power work
-
On 1/2, the UPS battery runtime estimate was very inaccurate and misled us about how much time we had for Facilities Management to complete the switchover.
-
On 1/3, the switchover to the rental generator was completed, but the UPS continued to drain. We powered off our storage right before losing power in an attempt to protect data. When the UPS attempted to power back on, it generated smoke. We powered it down, then placed it into bypass mode.
-
On 1/5, a Schneider engineer investigated the UPS issue. During that work, the rental generator powered off and we were restored to building power. We confirmed the rental generator was not the correct type for our power needs and that the UPS was not compatible with it.
-
After we recovered from the 1/5 outage and brought systems back online, we experienced transfer efficiency issues related to DNS.
-
First, our internal DNS service was in an abnormal state after the restart and was not functioning properly (e.g., nodes could not reach external repositories to install specific packages for testing).
-
Second, we needed to update DNS to include our newly rebuilt EL9 DTN servers.
-
After making the DNS updates and restarting the service, transfer efficiency returned to normal.
-
Campus Facilities Management successfully tested the building generator on 1/7, confirming it will automatically switch over in the event of another power outage.
-
We improved our process for rebuilding storage without migrating data to another system. We tested this and rebuilt a production storage server from EL7 to EL9. No data was lost. As a precaution, we also copied the data to an old storage server that will serve as a temporary storage area for future rebuilds.
-
Because one of our Frontier Squid servers (slate01) was not functioning properly and is not necessarily required, we disabled it and removed it from our site’s proxy list in CRIC.
OU:- Running smoothly
- Working on new xrootd servers, need Hiro's help
- Have new test SLURM jobmanager up, need to test cgroups v2 memory killing
-
-
11:00
→
11:10