US ATLAS Computing Facility
Facilities Team Google Drive Folder
Zoom information
Meeting ID: 996 1094 4232
Meeting password: 125
Invite link: https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09
-
-
13:00
→
13:05
WBS 2.3 Facility Management News 5mSpeakers: Alexei Klimentov (Brookhaven National Laboratory (US)), Dr Shawn Mc Kee (University of Michigan (US))
Working on discussion of USATLAS Ops https://docs.google.com/document/d/1pjwG1LAjOPWsSrdas4WvYYfoyn8NyLuk5u4L6opNb4Q/edit#heading=h.o4xh9deafqh6
Trying to finalize the HTC24 agenda, see https://docs.google.com/document/d/1em3rfH8lSwa5HnUomXwLKsSXFQGAV6S3pvT94ld-RY4/edit#heading=h.v2jvxcc92986
We plan to have a pre-scrubbing meeting on or shortly after June 20, 2024 with a goal of getting solid draft slides for the July scrubbing.
-
13:05
→
13:10
OSG-LHC 5mSpeakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
Software
-
Release next week
-
HTCondor-CE 23.8.0 in upcoming will contain breaking changes! It will deprecate the old job router config syntax! Expect an email with details this week
-
Backporting a few patches into XRootD
-
New timer expired error code https://github.com/xrootd/xrootd/issues/2264
-
Fix timing on throttle plugin https://github.com/xrootd/xrootd/pull/2262
-
Add config to allow deferring/disabling TLS auth https://github.com/xrootd/xrootd/pull/2269
-
Preview version of Kuantifier available! Matt sent an email to Armen and Eduardo with instructions
- OSG-LHC ARM server racked, cabled, and being built
-
-
13:10
→
13:30
WBS 2.3.1: Tier1 CenterConvener: Alexei Klimentov (Brookhaven National Laboratory (US))
- 13:10
- 13:15
-
13:20
Storage 5mSpeakers: Carlos Fernando Gamboa (Department of Physics-Brookhaven National Laboratory (BNL)-Unkno), Carlos Fernando Gamboa (Brookhaven National Laboratory (US))
-
Secure xroot support enabled at doors in dCache 9.2.17, including bug fix [www.dcache.org #10562]:Pools do not reload updated certificates.
-
Awaiting JBODs for new pool nodes (ETA 05/28) 14PB total to be deployed; Head Nodes delivered.
-
Rolling OS upgrade to RHEL 8 for warranted pool servers.
-
Using 2PB buffer for transparent data migration and OS upgrades; 15PB left.
-
Tuning DMZ pool activities ongoing.
-
MCTAPE write activity increased.
-
-
13:25
Tier1 Operations and Monitoring 5mSpeaker: Ivan Glushkov (University of Texas at Arlington (US))
- Added two EL9/Condor 21 CEs to OSG topology and CRIC
- Blacklisted (05/10) due to running out of file descriptors. This should be solved by EL9/cvmfs client update in the near future.
- Hi squid usage. “This is the new norm”
- BNL Tape firmware update: Allowed us to define and test blacklisting chain OSG/CRC for TAPE REST API.
-
13:30
→
13:40
WBS 2.3.2 Tier2 Centers
Updates on US Tier-2 centers
Conveners: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))- Reasonable running over the last 4 weeks when there was sufficient work.
- CERN did run out of work on two occasions causing significant draining
- Frontier error messages related to Varnish caused successful jobs to be marked a failed at AGLT2 & MWT2.
- NET2 ran pretty good with some minor draining incidents.
- Scheduled downtime today for power work..
- 3 more PB came online and is slowly filling.
- CPB managed to remove LSM but there was on knock-on effects.
- Ran into an XRootD issue caused by upper/lower case in web addresses that made transfers to sites running storm.
- They are working with Wei on getting the storage tokens enabled.
- Taiwan Tier 1 (TW-FTT) was supposed to be put online this week when network errors seemed to have been solved.
- Did not seem to happen. The site is still shown in test.
- Judith held a training session for Foreman / Puppet training for Aidan Rosberg (IU) and Zach Booth (CPB).
- Long discussion of coordinating cvmfs debugging for various problems seen at various sites.
- Kaushik was worried that the check that cvmfs is ok in the pilot wrapper before starting the pilot may cause issues.
- Reasonable running over the last 4 weeks when there was sufficient work.
-
13:40
→
13:50
WBS 2.3.3 Heterogenous Integration and Operations (HIOPS)
HIOPS
Convener: Rui Wang (Argonne National Laboratory (US))-
13:40
HPC Operations 5mSpeaker: Rui Wang (Argonne National Laboratory (US))
TACC
- HC jobs are running smoothly, brought the queue online today
- The queuing time is very long (2 hour job ~ few hours; 6 hour job ~1days; 10 hour job ~3-4days)
- Trying the job packing of 2 jobs x 3 nodes (6 hour) for worker
- upgrade atlas-cvmfsexec to 1.0.27 with cvmfsexec v4.39
Perlmutter
- 118K CPU hours added (total 818K)
-
13:45
Integration of Complex Workflows on Heterogeneous Resources 5mSpeaker: Doug Benjamin (Brookhaven National Laboratory (US))
-
13:40
-
13:50
→
14:10
WBS 2.3.4 Analysis FacilitiesConveners: Ofer Rind (Brookhaven National Laboratory), Wei Yang (SLAC National Accelerator Laboratory (US))
-
13:50
Analysis Facilities - BNL 5mSpeaker: Dr Quilan Huang (BNL)
-
13:55
Analysis Facilities - SLAC 5mSpeaker: Wei Yang (SLAC National Accelerator Laboratory (US))
-
14:00
Analysis Facilities - Chicago 5mSpeaker: Fengping Hu (University of Chicago (US))
- starting development on OpenAI assistant for AF.
- need to collect documentation and discourse conversations
- adding GPU info and condor queue info to AF monitoring
- Highlights of AF maintenence done on Monday/Tuesday
-
Migration of all servers to Enterprise Linux 9
-
HTCondor updated to the OSG 23 release series
-
Update of the /data filesystem to Ceph v17
-
System BIOS and firmware updates on all servers
-
- starting development on OpenAI assistant for AF.
-
13:50
-
14:10
→
14:25
WBS 2.3.5 Continuous OperationsConvener: Ofer Rind (Brookhaven National Laboratory)
- DC24 Report will be submitted next week, so please consider giving it a once-over if you haven't already (link)
-
14:10
ADC Operations, US Cloud Operations: Site Issues, Tickets & ADC Ops News 5mSpeaker: Ivan Glushkov (University of Texas at Arlington (US))
- ADC:
- ~A week of low GRID occupancy due to lack of simulation in the system
- Rucio:
- xrootd to storm sites - transfers fail due to xrootd bug. It will be fixed in version 5.7 this summer. Temporary patch is removing distances between these sites.
- HC:
- Short mass-blaklisting earlier today due to a Panda bug.
- CVMFS Monitoring - available, but consists of two separate categories of errors - not avle to get pilot and not being able to access the cvmfs
- Started Meetings:
- First ADC Fabrics Meeting (Indico:1414901)
- “to facilitate communication and contributions between ADC and infrastructure providers”. Monthly
- US ATLAS Distributed Computing Ops (Doc)
- Same idea as the ADC Ops meeting.
- Daily, 9:30 AM CDT / 10:30 AM EDT / 4:30 PM CEST / 10:30 PM CST
- Started this Monday with trail period of one week.
- Already addressed: SWT2 Storage Tokens, MWT2 CA related transfer errors, monitoring for CVMFS issues, etc.
- First ADC Fabrics Meeting (Indico:1414901)
- ADC:
-
14:15
Services DevOps 5mSpeaker: Ilija Vukotic (University of Chicago (US))
- XCache & VP
- MWT2 xcaches back in operation. Reactivated VP queue today
- Still having 8 xcaches serving AF
- VP working fine
- Varnishes
- added an instance in NRP for NET2. Works fine.
- will try configuring DNS Anycast in Cloudflare so all the varnishes would be behind the same name. Failovers would be automatic. 5$/node/month.
- ServiceX
- Production instance on AF restarted. Still running with a lot of manual fixes.
- Today moving all the testing instances to River-dev cluster
- XCache & VP
-
14:20
Facility R&D 5mSpeaker: Lincoln Bryant (University of Chicago (US))
- Successful Kubernetes tutorial and workshop last month, a lot of interesting ground covered
- 20+ successful single-node K8S clusters built with Kubespray
- Kueue (multi-user fair share scheduling), Karmada (multi-cluster), stretched K8S over Wireguard VPN, Lens, Anycast DNS, GRACC/KAPEL, etc
- Stretched BinderHub platform deployed across UM, MSU, UC, IU, UVic: https://rp1.hl-lhc.io/
- Want to work with NET2 and SWT2 to meet our milestone of integrating all T2s
- Bringing Aidan Rosberg, new hire at IU, up to speed
- Aidan is working on rebuilding several nodes at IU and learning Kubernetes in the process
- Reana installed on UC AF, working on adding ATLAS IAM auth
- KAPEL/GRACC integration ongoing
- Ongoing work to push identity information down into Binder/Jupyter containers - such that we can have POSIX identity + provision storage, etc.
- Successful Kubernetes tutorial and workshop last month, a lot of interesting ground covered
-
14:25
→
14:35
AOB 10m
-
13:00
→
13:05