US ATLAS Tier 2 Technical
Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.
-
-
11:00
→
11:10
Introduction 10mSpeakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
Quick links:
General news
- Quarterly reports are overdue. Please, if you haven't yet, submit it as soon as possible.
Upcoming meetings:
- WLCG/HSF Workshop 2026 [Nov 2nd -- 6th, Bologna]
- 57th LHCOPN-LHCONE meeting [Oct 1st -- 2nd, Naples]
Open tickets:
- ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
Operations:
- Site production during the previous 2 weeks: AGLT2, MWT2, NET2, SWT2 (CPB, OU), TW
- TW
%20(5).png)
AGLT2
.png)
MWT2
%20(1).png)
NET2
%20(2).png)
SWT2/CPB
%20(3).png)
SWT2/OU
%20(4).png)
- 11:10 → 11:20
-
11:20
→
11:30
AGLT2 10mSpeakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
1) updated condor from 25.0.9 to 25.0.12 to address the security issues
2) updated dcache from 11.2.4->11.2.5, the update last about 1hour, and AGLT2 was set offline for half an hour. No downtime was declared.
3) Updated and rebooted all nodes but the UM work nodes (due to Lustre version limit) to the latest 5.14.0-687.25 and 26 kernel to address several CVE fixes and the latest firmware when they apply
4) plan to swap a Rack in the UM Tier2 room, this Rack hosts about 12% of the AGLT2 data, we set them as read only 4 days in advance, and not declare downtime. The work is estimated to last a whole day.
-
11:30
→
11:40
MWT2 10mSpeakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
-
11:40
→
11:50
NET2 10mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
Following up on intermittent issues with efficiency, with multiple causes. One is pods using more storage than they are allocated. We are looking into allowing pods to temporarily exceed this limit, since most pods don't come close to it and nodes generally run with at least 60% free storage. If this doesn't work, we may have to increase the amount of storage per core in CRIC and see if the efficiency improvement outweighs the loss of slots. Spikes of backoff errors may also be due to issues with diskIO overload, we are looking into this as well.
Kuantifier is now running on the cluster. We lost some slots while it was being implemented because machines were being used for development and troubleshooting.
-
11:50
→
12:00
SWT2 10mSpeakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
SWT2_CPB:
-
We implemented a second XRootD redirector for redundancy and rebuilt our EL7 XRootD redirector to EL9.
-
We performed this in stages: Tested this in the test cluster, implemented the second EL9 redirector, removed the EL7 redirector, transitioned the EL7 redirector to EL9, then added it back into production.
-
We had a configuration error that led to errors and HC blacklist on 7/10. We realized that we needed to restart the services on storage servers and that the configuration on one of the XRootD redirectors was incorrect. We restarted the XRootD service on these nodes and corrected the configuration.
-
Nodes in SWT2_CPB_K8S nodes are using EL7 admin node for DNS lookups. Because the newer version of XRootD uses hostname instead of IP address for storage and EL7 admin did not have EL9 storage in DNS, jobs failed on these worker nodes on 7/15. We changed DNS on admin to fix this issue.
-
EL7 admin node updated DNS changes on 7/20, removing the fix we performed on 7/15, which started causing the same errors as before. We fixed DNS again and added a more permanent solution. We are discussing the SWT2_CPB_K8S queue.
-
Once these issues have been fixed, we have not seen issues.
-
GGUS-Ticket-ID: #1003053 - IGTF CRLs & fetch-crl SHA1 signature validation (SWT2_CPB)
-
We removed gk05 from the DNS round robin.
-
We rebuilt gk05 and are using it for testing of other tasks, such as testing a potential network upgrade.
-
We plan to inspect this server more thoroughly and will add it back to the DNS round robin to test for the same errors.
OU:- GPU nodes updated to CUDA 13.3
- Working on migrating storage, currently waiting for Petr or Fabio to see why our gfal-copy transfers succeed, while rucio's tests fail
- Still some fraction of condor-hold failures, not sure why; no indication of any local failures
-
-
11:00
→
11:10