US ATLAS Tier 2 Technical
Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.
-
-
11:00
→
11:10
Introduction 10mSpeakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
News:
- Capability mini-challenges ongoing this week!
- If you have any updates, feel free to share during the round table (and enter in the minutes)
- New dCache version 11.2 with fireflies released.
- If your site is testing it and has results, please share news during the rount table (and enter in the minutes)
- Keep the capacity and services spreadsheets updated. Keep CRIC and OSG topology updated when servers are added or retired.
- Next procurement meeting this Friday (2/6), at 11am Eastern (indico:1647183)
Upcoming meetings:
- Just finished: WLCG Open Technical Forum (OTF) #8 on tape evolution, slides available including our experience with Tier-2 tapes in US ATLAS
- Next week: ATLAS S&C meeting [Feb 9-13 at CERN]
-
CHEP 2026 [23-29 May in Bangkok, Thailand]
- LHCOPN-LHCONE meeting #56 [Apr 15-16 in Montreal]
- HEPiX Spring 2026 Workshop [Apr 20-24 in Lisbon]
- dCache workshop [May 6-7 at NIKEFF]
Open tickets:
- ggus:1001722 NET2: transfer error
- ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
- ggus:3559 SWT2/OU: Dual-stack [on hold]
- ggus:1001633 SWT2/CPB: all transfers to IFIC-LCG2 and TECHNION-HEP failing due to SSLHandshakeException
- ggus:1001382 TW-FTT: failing transfers as SOURCE due to certificate issue
Operations:
- AGLT2
- MWT2
- NET2
- SWT2/CPB
- SWT2/OU
- Capability mini-challenges ongoing this week!
-
11:10
→
11:20
TW-FTT 10mSpeakers: Eric Yen, Felix.hung-te Lee (Academia Sinica (TW)), Yi-Ru Chen (Academia Sinica (TW))
- Several worker nodes were added, increasing the total to 3,120 slots.
- Generally, site is running smoothly.
- ggus:1001382 : No updates at this time.
-
11:20
→
11:30
AGLT2 10mSpeakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
Testing removing IPv4 on LHCONE started this morning
Short test successful this morning
with both campuses removing their IPv4 announcements
can still reach places over IPv4 via commodity internet
Now starting a 2 day test with ESnet removing IPv4
AGLT2 NET2 for ATLAS and Purdue UNL for CMS
Mini data challenge for testing host tuning (fasterdata)
Hiro orchestrating writing and reading AGLT2 <> MWT2
Tested baseline on Monday
Noticed our dcache system was NOT redirecting on read
This explains the poor read performance in previous tests (plus other observations)
Tested tuned configuration on Tuesday
First without correcting redirection
Again after correcting dcache config
But MWT2 storage seemed saturated with merging requests
Will need to repeat this test
Added gate03 to A/R monitoring (only had gate01 before)
Needed to set in_report="True" for gate03 in CRIC
Also found and corrected issue with periods of UNK status
Test jobs were idle, waiting and timing out
We had a low-ish quota on the number of test jobs
and a configuration issue was leaking other jobs into that quota
All corrected now.
We will need to request a correction for January A/R
Earlier mini data challenge tests had identified some pool nodes with poorer IO performance
Believed to primarily come from hardware raid6 stripe size; default 64k; our choice/compromise 512k
Started a campaign of reconfiguring those nodes; 2 done already; this requires draining. -
11:30
→
11:40
MWT2 10mSpeakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
- Brief network issue at UChicago that took MWT2 offline for a couple hours (1250 - 1450 Chicago time) on Feb 3
- Testing dcache 11.2 golden release. Looking at scheduling the downtime to upgrade soon - Asking about supported java version
- Mini capacity challenge Feb 2-3: Feb 2 without tunings and Feb 3 with. May need to rerun with tunings due to job overload and the UChicago network outage on the 3rd
- Discussing potential ancillary purchases (head nodes, spares, etc.) with this cycle before PI meeting on Friday
-
11:40
→
11:50
NET2 10mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
1001545: Updated our CA packages so that transfers to EELA are no longer failing
1001722: Associated with last night's blacklisting. The issue seems to be with one of our pools, which was also one of the problem pools on Monday night. It seems that once enough transfers accumulate on this pool, its throughput drops to practically nothing, leading to the failures observed both last night and the night before. We are investigating.
Capability challenges:
- Host tuning: We had to postpone. Our team is working on updating tape hardware at the moment and can't assist on storage tests.
- IPv4 blackout: "short test" happened without problems. Long test starting now.
- SENSE: Testing data transfers between UMass and UChicago.
-
11:50
→
12:00
SWT2 10mSpeakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
SWT2_CPB:
-
We finished migrating data from all of our old storage nodes and officially retired them.
-
This includes eight MD3460 storage arrays.
-
We will continue to use them for various purposes such as testing or any other way they can support our processes, but they are no longer used as part of our XRootD storage.
- 3.025 PB was retired.
-
We had one problematic storage server with a failed OS drive and data drive. The OS drive was one of two, so the server continued to run without disruption.
-
We created a data backup before coordinating with the vendor for warranty-covered repairs.
-
Both its failed OS drive and data drive have been replaced and are working normally now.
-
We finished checking worker nodes to identify which nodes has reached end of life after the power outages and repaired other worker nodes to put them back into service. About 21 worker nodes were taken out of service.
-
We added IPv6 to our Squid servers, but are waiting for DNS change by campus networking for fsquid.atlas-swt2.org.
-
We noticed a UTA_SWT2-Squid service entry in CRIC, and are looking into whether this should be removed for CRIC cleanup.
-
We have created an EL9 XRootD redirector and Squid instance in our test cluster.
-
We are continuing to develop the EL9 Puppet modules before implementing them into the production cluster.
-
GGUS-Ticket-ID: #1001672
-
Our Squid service status stopped appearing in monitoring on January 28.
-
We found that this object was intentionally disabled in CRIC by cloud ops (not SWT2).
-
After discussion with ADC Coordinator, we re-enabled this object.
-
GGUS-Ticket-ID: #1001633
-
We are continuing to experience transfer errors with SWT2_CPB as the source site to IFIC-LCG2 and TECHNION-HEP.
-
We have been investigating and are not seeing any issue at our site so far.
-
There are other source sites affected and GGUS tickets created related to this.
-
We did experience some failed jobs related to this due to timeouts over the past two weeks.
-
We need help from experts to assist IFIC-LCG2 and TECHNION-HEP to find a solution.
-
-
11:00
→
11:10