US ATLAS Computing Facility
Facilities Team Google Drive Folder
Zoom information
Meeting ID: 996 1094 4232
Meeting password: 125
Invite link: https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09
-
-
13:00
→
13:10
WBS 2.3 Facility Management News 10mSpeakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn Mc Kee (University of Michigan (US))
-
13:10
→
13:20
OSG-LHC 10mSpeakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
Site admin office hours Tue Jan 25 1-4pm Central. Register here: https://docs.google.com/forms/d/e/1FAIpQLSdnvnv3uFdKN5MiVFmpFfsIYaZVZDLbpJUvTBprBsGpsSgKxQ/viewform
Release
- HTCondor-CE 5.1.3 (3.5 upcoming + 3.6)
- CVMFS 2.9.0 (3.6 only)
- CA cert updates (3.5 + 3.6)
- HTCondor 9.5.0 (3.6 upcoming)
- HTCondor 9.0.9 (3.5 upcoming + 3.6)
Token Transition
- osg-scitokens-mapfile-4-1 contains default ATLAS token -> local user mappings, please update!
- If your site only supports ATLAS, GLOW, and OSG: you can update to OSG 3.6! https://opensciencegrid.org/docs/release/updating-to-osg-36/
- CEs on token-supporting versions of HTCondor-CE (let me know if you expect to see your CE here but it isn't listed!)
- bgk01.sdcc.bnl.gov
- bgk02.sdcc.bnl.gov
- gate01.aglt2.org
- gate02.grid.umich.edu
- gate04.aglt2.org
- gpce03.fnal.gov
- gpce04.fnal.gov
- gridgk01.racf.bnl.gov
- gridgk02.racf.bnl.gov
- gridgk03.racf.bnl.gov
- gridgk04.racf.bnl.gov
- gridgk06.racf.bnl.gov
- gridgk07.racf.bnl.gov
- gridgk08.racf.bnl.gov
- iut2-gk.mwt2.org
- osg-gk.mwt2.org
- spce01.sdcc.bnl.gov
- spce02.sdcc.bnl.gov
- uct2-gk.mwt2.org
- CEs on old versions of HTCondor-CE
- atlas-ce.bu.edu
- gk01.atlas-swt2.org
- gk04.swt2.uta.edu
- grid1.oscer.ou.edu
- mwt2-gk.campuscluster.illinois.edu
-
13:20
→
13:50
Topical ReportsConvener: Robert William Gardner Jr (University of Chicago (US))
-
13:20
Facility Readiness for Run 3 30m
-
13:20
-
13:50
→
13:55
WBS 2.3.1 Tier1 Center 5mSpeakers: Doug Benjamin (Brookhaven National Laboratory (US)), Eric Christian Lancon (Brookhaven National Laboratory (US))
-
13:55
→
14:15
WBS 2.3.2 Tier2 Centers
Updates on US Tier-2 centers
Convener: Fred Luehring (Indiana University (US))- It was a very good two weeks.
- The generral draining on 1/13 was caused by two of the aipanda VMs becoming overloaded which resulted in HC offlining most grid sites. A user was uploading 500 MB tarball which caused the aipanda VMs to crash. The ADC team banned the user and restarted the affected VMs.
- The draining of MWT2 today is for the quarterly scheduled preventative maintenance at the Illinois site.

- I suspect at this point we won't need to further discussion of the site's readiness for run 3 because of the prior talk.
- Get those purchases in!
-
13:55
AGLT2 5mSpeakers: Philippe Laurens (Michigan State University (US)), Dr Shawn Mc Kee (University of Michigan (US)), Prof. Wenjing Wu (University of Michigan)
Smooth running overall.
1/5/2021
One of the data switches in the UM Tier3 room (sw9-d-01) got stuck,
and 6 work nodexs which are connected to the switch lost connection.
The solution is to power cycle the switch.1/17/2022
At 18:11, one dcache pool node (umfs06) rebooted by itself (not clear why).
dcache was not restarted until 19 hours later, manually,
which caused 500 jobs failure with stage-in errors.Currently finalizing quotes for end of cycle purchase.
- R740xD2 with 24x 18TB disks
- R6525 with AMD 7413 -
14:00
MWT2 5mSpeakers: David Jordan (University of Chicago (US)), Jess Haney (Univ. Illinois at Urbana Champaign (US)), Judith Lorraine Stephen (University of Chicago (US))
UC:
- Trouble with two dcache nodes over the past couple weeks. Both are back up currently, but led to some job failures. Suspect that some hardware is bad on one. Will investigate when it's drained.
- Second physical hardware move to new data center next week.
- New compute and storage to arrive in the next few weeks.
- Updated MWT2-TEST to osg 3.6
IU
- Updating compute nodes to condor 9
- Machines moved to IU from UC are cabled and will start to get put back in production soon.
UIUC
- ICC PM today, site offline.
-
14:05
NET2 5mSpeaker: Prof. Saul Youssef
1. In the process of retiring 2.5 racks of 3TB storage (770 TB useable).
2. Added 4 nodes to GPFS xrootd cluster
3. 1&2 greatly improved staging performance, GPFS slowness issues
4. Remaining hardware orders finalizing through BU purchasing... 10 new transfer nodes and 3.8 PB NESE Ceph storage being purchased. No new worker nodes.
5. 88 worker nodes schedule to arrive from DELL March 3
6. Lining up collaborations and organization for bare metal cluster & UMass expansion.
7. Preparing to upgrade NET2-NESE networking to 400Gb/s.
8. NESE Tape commissioning continues with NESE team, Xin, Alexei and ADC.
Smooth operations in the past 2 weeks.
-
14:10
SWT2 5mSpeakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))
SWT2_CPB -
- Added second host to the webdav pool, plus upgraded XRootD to v5.4.0. Better performance and stability since implementing these changes.
- Working with Hiro to optimize the concurrency settings in FTS.
- Odd routing for transfers to RAL - not using LHCONE path?
- Submitted quotes to our procurement office for the latest hardware purchase: 3 PB of storage, 4608 logical job slots in WN's.
UTA_SWT2 -
- Working with personnel at the data center to finalize our shutdown & hardware move dates.
OU:
- Not much to report, running stably.
- Still occasional xrootd hangups, hopefully upgrading backend storage to 5.4.x will fix that.
- Working with OU Purchasing to put Dell quote through, to spend remaining hardware funds.
- It was a very good two weeks.
-
14:15
→
14:20
WBS 2.3.3 HPC Operations 5mSpeakers: Lincoln Bryant (University of Chicago (US)), Rui Wang (Argonne National Laboratory (US))
-
14:20
→
14:35
WBS 2.3.4 Analysis FacilitiesConvener: Wei Yang (SLAC National Accelerator Laboratory (US))
- 14:20
-
14:25
Analysis Facilities - SLAC 5mSpeaker: Wei Yang (SLAC National Accelerator Laboratory (US))
-
14:30
Analysis Facilities - Chicago 5mSpeakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
-
14:35
→
14:55
WBS 2.3.5 Continuous OperationsConvener: Ofer Rind (Brookhaven National Laboratory)
- Token readiness discussion at last week's WFMS weekly meeting
- Follow up discussions at S&C week
- Ongoing transfer issues at CPB (RAL IPV4 issue has been resolved but backlog still a problem)
- Pilot 3 being deployed
- Evaluating Run 3 readiness (see above)
-
14:35
US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5mSpeakers: Mark Sosebee (University of Texas at Arlington (US)), Xin Zhao (Brookhaven National Laboratory (US))
-
14:40
Service Development & Deployment 5mSpeakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
Analytics
- additional ES nodes drained for transport
- three new logstash based ingresses
- all the ML platform images are being upgraded today
XCaches
- fixes to gStream reports
- testing ephemeral storage changes
- some unexpected restarts due to k8s liveness probes failing.
VP
- running fine
ServiceX
- some developments got merged
- needs more testing
-
14:45
Kubernetes R&D at UTA 5mSpeaker: Armen Vartapetian (University of Texas at Arlington (US))
- Token readiness discussion at last week's WFMS weekly meeting
-
14:55
→
15:05
AOB 10m
-
13:00
→
13:10