US ATLAS Tier 2 Technical
Meeting to discuss technical issues at the US ATLAS Tier 2 site. The primary audience is the US Tier 2 site administrators but anyone interested is welcome to attend.
-
-
11:00
→
11:10
Introduction 10mSpeakers: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
Quick links:
General news
- Ongoing discussion about equipment funding based on new end-of-CA estimate ongoing.
- Still facing serious supply chain problems (several vendors refusing purchase orders)
- In discussions with management about the situation.
Upcoming meetings:
- 57th LHCOPN-LHCONE meeting [Oct 1st -- 2nd, Naples]
- #85 ATLAS S&C week [Oct 5th -- 9th, CERN]
- HEPiX Fall 2026 Workshop [Oct 19th -- 23rd, Nebraska] -- Possible USATLAS/USCMS facility discussion... to be announced.
- WLCG/HSF Workshop 2026 [Nov 2nd -- 6th, Bologna]
Open tickets:
- ggus:1001568 SWT2/OU: xrootd version higher than 5.7.0 needed
- ggus:1003053 SWT2/CPB: IGTF CRLs & fetch-crl SHA1 signature validation
- 11:10 → 11:20
-
11:20
→
11:30
AGLT2 10mSpeakers: Daniel Hayden (Michigan State University (US)), Philippe Laurens (Michigan State University (US)), Shawn Mc Kee (University of Michigan (US)), Dr Wendy Wu (University of Michigan)
Starting to test RHEL10
built a test pool node VM to work on ansible updates
Plan to have some fraction of the newer dCache pool nodes on RHEL10
for comparison during next capacity challenge
built a test WN VM, it finished 18 ATLAS jobs and 3 failed.S3 storage problem
Certificate expired on Aug 7th,
Replaced it with an IGTF cert we had pre-emptively renewed before InCommon transitioning to CERTinext,
Still working on access to new IGTF/OV cert from InCommon/CERTInext -
11:30
→
11:40
MWT2 10mSpeakers: Aidan Rosberg (Indiana University (US)), David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Fengping Hu (University of Chicago (US)), Fred Luehring (Indiana University (US)), Judith Lorraine Stephen (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
-
11:40
→
11:50
NET2 10mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
Blacklisted yesterday with everybody else. There was also a second dip in the number of running slots around 4:30 am Eastern this morning, shared by all US sites, do we know what caused that?
Making progress on limiting backoff errors, traced the nvme errors we were seeing to an issue with the kernel, which we have a workaround for. Additionally, the newer OKD version enabled more compression of resources under load by default, meaning that the nodes could run at 15-20% over CPU capacity. This turned out to be too much, so we turned that off. Over the last couple of days the number of backoff errors seen is much reduced. We are still running at a lower amount of slots because of issues with the two machines that required disk replacements.
The tape is out of downtime, we are testing to see if the throughput has improved with the fixes that were made.
-
11:50
→
12:00
SWT2 10mSpeakers: Andrey Zarochentsev (University of Texas at Arlington (US)), Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Zachary Thomas Booth (University of Texas at Arlington (US))
SWT2_CPB:
-
Investigated reported issues concerning deletions by a user.
-
This may be an issue involving DDM, but discussions are ongoing.
-
We are working toward understanding this issue better.
-
jemalloc has significantly helped with memory management on our XRootD Proxy servers.
-
Two of our XRootD Proxy servers still restart their service from OOM events from time to time, but less frequently than before using jemalloc.
-
The changes we made to the service to help protect against significant disruptions is working as intended. Restarts are quick, minimizing transfer disruptions during OOM events and failed service.
-
Removed one of our XRootD Proxy servers temporarily to investigate hardware issues and other possible causes of memory issues.
-
We have added this server back into service.
-
Testing new configuration for perfSONAR machines with current plans to rebuild our perfSONAR machines with this new configuration for improved security.
-
New planned build is testpoint with containers.
- Thanks for the quick patch from perfSONAR developers to the security issue identified at SWT2.
- The developers had additional suggestions for all US ATLAS sites - we should follow up in a dedicated meeting.
-
Campus networking has performed a network upgrade in the past two weeks.
-
The upgrade doubled the bandwidth of the bundle connecting the machine room from 40 to 80 Gbps. This is part of a multi-step phased networking upgrade over the next year.
-
We are still monitoring to confirm this change.
-
They attempted the upgrade on 8/21. There were issues with transfers. They responded quickly by disabling the new connection to investigate.
-
They found the issue, resolved it, then added the new connection on 8/29. We monitored and there were no issues.
-
Our problematic CRAC unit was repaired and powered back on. We monitored then put our drained worker nodes back into service on 8/24 and 8/25.
-
Previous issues:
-
GGUS-Ticket-ID: #1003053 - IGTF CRLs & fetch-crl SHA1 signature validation
-
Problem with connect between one of our DTNs and BEIJING site (check crl ) (we have not made any changes the past two weeks, still investigating, but we remember this issue)
-
We have a few jobs running on the TEST-CE. We do not have hammercloud jobs on this cluster.
-
From time to time we see the "Unknown" status of our CE. We do not understand reason now
OU:- Had a failed raid6 drive and some xfs file system corruption on one of the 7 old xrootd data servers. Fixed.
- Migration to new storage is making progress. Getting up to 4 GB/s inbound WAN transfers to new storage now.
- Waiting to hear back from Timo et al about switching Panda over to start using the new storage for stage-in/out.
-
-
11:00
→
11:10