US ATLAS Computing Facility
Facilities Team Google Drive Folder
Zoom information
Meeting ID: 996 1094 4232
Meeting password: 125
Invite link: https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09
-
-
13:00
→
13:10
WBS 2.3 Facility Management News 10mSpeakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn Mc Kee (University of Michigan (US))
We need to finish the updates for the milestones: https://docs.google.com/spreadsheets/d/1DsMH-16v7bJy6qEkvTEWdLCkfeAXC6VpL019rUpcET8/edit#gid=533220483
Prescrubbing/technical meeting Bloomington, in June 5-8 https://indico.cern.ch/event/1273590/
CHEP 2023 https://www.jlab.org/conference/CHEP2023 coming up in May
ATLAS DDM meeting today https://indico.cern.ch/event/1276964/ (Was recorded, see attached video/transcript)
-
13:10
→
13:20
OSG-LHC 10mSpeakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
Release
- XRootD 5.5.4 with xrdcl-http and frontier-squid 5.8.2 are available in osg-testing
- HTCondor 9 EOL on May 11, we're expecting to release HTCondor 10.0 in OSG 3.6 release before then
- There are known issues with scitokens-cpp-1.0.0 from EPEL. If your 'condor' user has a HOME directory then you are unaffected but otherwise, avoid this version
-
13:20
→
13:40
WBS 2.3.5 Continuous OperationsConvener: Ofer Rind (Brookhaven National Laboratory)
- Massive outage over the weekend caused by expired intermediate CERN CA cert
- May affect user access, e.g. to GGUS - may need to remove CERT from your browser
- WT2 is accepting jobs again through ARC-CE (GGUS); perhaps not up to full scale running yet
-
13:20
US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5mSpeaker: Mark Sosebee (University of Texas at Arlington (US))
-
13:25
Service Development & Deployment 5mSpeakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
XCache
- working fine
- Andy debugging issues with Oxford
- RHUL now going storageless
- will upgrade to 5.5.4 probably before end of the week
VP
- works fine
- rucio integration - waiting on one more PR to get merged
- need to find time to follow up with BNL_VP
Varnish
- working fine at both MWT2 and AGLT2
ServiceX
- now AF deploys new unified ServiceX in addition to Uproot and xAOD versions.
- deployed it at FAB nodes at CERN
- this is IPv6 only cluster, all kinds of issues cropped up.
Elasticsearch
- Upgrade to 8.7.0 probably tomorrow
CREST
- Storage optimization
- More metrics monitoring
- More logging improvements
- need TLS on loadbalancer
-
13:30
Kubernetes R&D at UTA 5mSpeaker: Armen Vartapetian (University of Texas at Arlington (US))
- Current production cluster SWT2_CPB_K8S :
- Is running fine. The 300s timeouts on rucio get, a known problem, fixed in the new rucio release. Haven't seen those errors recently.
- Queue started to drain late Saturday (other K8S queues as well). Noticed problem on the Harvester side with aipanda169 (k8s instance) down. Contacted Harvester support - they applied a fix which resolved it.
- New Test cluster inside the SWT2_CPB main cluster network :
- The new cluster with a master node and 4 worker nodes as a starting point.
- K8S installation went mostly without issues, some hiccup with pod network, but was resolved.
- Some delay with the Harvester setup software, not working with K8S recent releases. While waiting for response from harvester support managed to find a solution which worked with harvester configuration.
- The SWT2_CPB_K8S_TEST queue was created. Jobs were not coming. Appeared to be a HC issue (https://its.cern.ch/jira/browse/HC-1372), they fixed it. Though it was down a couple of more times during the past week.
- Working fine on the K8S side, but issues with copytools.
- Today Patrick managed to make LSM work for SWT2_CPB_K8S_TEST, and HC test jobs now are finishing fine.
- We have Ganglia for monitoring. Installed Prometheus, working on configuration.
- Some issues with Disk Pressure on the master node. We need to increase the size of the main partition, but Rocks will need to reinstall the node...
- Current production cluster SWT2_CPB_K8S :
- Massive outage over the weekend caused by expired intermediate CERN CA cert
- 13:40 → 13:45
-
13:45
→
14:05
WBS 2.3.2 Tier2 Centers
Updates on US Tier-2 centers
Convener: Fred Luehring (Indiana University (US))- Reasonable running over the past 30 days,
- The plot shows 2 large draining incidents: 4/10-11 where there weren't enough jobs available and 4/22 where an expired CERN root certificate. caused a huge disruption.
- BU_NESE was affected by the expired certificate and started failing transfers. We need to figure out what to do in this case since BU is not managing the end point any longer.
- The quarterly reporting was completed over last weekend.
- Progress is good for NET2.1 but I will let Rafael/Eduardo describe that.
- And a new AlmaLinux9 gatekeeper went online at OU as did the new compute servers ordered last year.
- I am working with WFMS group on getting Panda queue parameters to describe the distribution of memory/slot so an appropriate job mix is assigned to the sites.
- Please let me know the status of your FY23 procurement.
-
13:45
AGLT2 5mSpeakers: Philippe Laurens (Michigan State University (US)), Dr Shawn Mc Kee (University of Michigan (US)), Prof. Wenjing Dronen
Incidents:
We had a report that “Some event index jobs are failing because they can not access the AGLT2 data ''. The reported files seem to be accessible but we found msufs15 had crashed. dCache was restarted to solve this issue.
One of the dCache pools msufs02 rebooted itself during the night, and dCache service was not started for about 5 hours, which resulted in 27% of the jobs failing at accessing the files. Like for msufs15 earlier, we don't see anything curious in log/messages.
No hardware error either in OMSA alertlog, nor in esmlog. Later the same day dcache pool msufs02_10 had disabled itself with a file system complaint. Ran xfs_repair and started the pool. No clear explanation; this may have been some residual file system damage from the crash and reboot.
2023 hard purchase:
The UM received 14 storage nodes and 5 work nodes, but 2 storage nodes were damaged during shipping. Dell has placed a new order for 2 replacing nodes, estimated to arrive at the end of May. We brought online 7 dCache nodes, but because the ipv6 setting missed some configuration from cfengine, the nodes had auto created IPV6 addresses and did not have the routing rules to the MSU nodes for IPV6, and this cause 50% of the job failure. We manually fixed the IPV6 issues in order to have a quick recovery. One of the new nodes had its network stopped working after rebooting, this was caused by a mislocation of the breakout cable. We recovered this node by disabling one of the ports of the LCAP bonding.
MSU has received all 2023 equipment ordered.
6x R6525 with AMD 7443 and 4x R740xd2 with 20TB drives being racked today. -
13:50
MWT2 5mSpeakers: David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Judith Lorraine Stephen (University of Chicago (US))
- Storage node at UC had a failed disk that caused some file index issues. Still working on getting a replacement disk, but the machine was rebooted and seems to be working properly now.
- Updated condor to 10 at all three sites. Waiting on condor-ce until 6 is out of testing.
- Two new gatekeeper to serve UIUC compute first to help with refilling and load balance.
- WLCG network monitoring milestone complete for UC and IU.
- GGUS 161670: Seems to be an upstream issue with CRIC and not a site admin issue.
- Will be updating elasticsearch to the latest version tomorrow, April 27.
- Most of the IU compute nodes have arrived and are being racked.
-
13:55
NET2 5mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
- Transfer of the machines is finalized. Nothing left with BU.
- All physical fiber connections at MGHPCC done.
- Basic management network established (for access to routers, switches, PDUs, ...)
- BGP with ESnet/LHCONE (via NoX) established and being debugged.
- Some weird routing behavior, but we are trying to debug
- Dedicated connection between the MGHPCC data center and UMass physics department (for management and monitoring) established and being debugged.
- Configuration of BGP with NESE ongoing. We expect it to be ready soon and then we will start debugging it.
- Quotes for racks with RDHX received, working on a schedule with university.
- From now until the RDHX racks are available, the machines will operated on a separate pod with in-row cooler which may be able to absorb the full power since it is completely empty.
- Monitoring for dCache and network ongoing.
- Effort to generate SRR for dCache. Final configuration being discussed on MM.
- Installation of perfSonar node and SNMP monitoring ongoing.
-
14:00
SWT2 5mSpeakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))
UTA
- Internal Network upgrade completed
- All compute nodes rebuilt during process
- Working on K8s
- Compute node build procedure OK
- Master node build may need adjustment
- Using LSM for stage-out, slightly modified for use by K8s jobs
- Next Up deploy compute nodes and balance power
OU
- Remaining compute nodes deployed, running well
- Upgraded CE OSG and O/S last week, but upgrade of scitokens-cpp from 0.7.3 to 1.0.0 broke HTCondor-CE submission
- Downgrading back to 0.7.3 made it work again
- Condor home directory reconfiguration may fix this issue, will investigate
- Internal Network upgrade completed
- Reasonable running over the past 30 days,
-
14:05
→
14:10
WBS 2.3.3 HPC Operations 5mSpeakers: Lincoln Bryant (University of Chicago (US)), Rui Wang (Argonne National Laboratory (US))
- Starting to scale up at TACC, ~1700 jobs completed in the last 24h with 93% efficiency on avg.
- ~450 jobs failed, mostly due to being cancelled by Harvester
- harvester sweeper plugin will kill the entire slurm job if a single payload gets closed/reassigned by JEDI. Needs followup. Potentially 1 killed PanDA job causes 34 other other jobs to be collateral damage.
- Needs followup with WFMS?
- ~450 jobs failed, mostly due to being cancelled by Harvester
- Perlmutter running some jobs after acceptance testing, ~2100 jobs finished in the last 24h with 90% efficiency on avg
- ~200 jobs cancelled (similar to TACC?),
- ~460 jobs failed due to " failed before adding files : GUID is inconsistent between jobReport and pilot report for HITS.33117584._009047.pool.root.1"
- needs followup
- Down again for maintenance today through tomorrow.
- Working through the latest CSV failed log file xfers from BNLHPC disk and matching it up with jobs in the Globus logs.
- Working on testing GPU jobs in the 'production' instance of Harvester @ Perlmutter.
- Jobs being evicted due to OOM for reasons we don't yet understand
- Keeping all configuration here: https://github.com/usatlas/harvester-config-perlmutter
- Starting to scale up at TACC, ~1700 jobs completed in the last 24h with 93% efficiency on avg.
-
14:10
→
14:25
WBS 2.3.4 Analysis FacilitiesConveners: Ofer Rind (Brookhaven National Laboratory), Wei Yang (SLAC National Accelerator Laboratory (US))
-
14:10
Analysis Facilities - BNL 5mSpeaker: Ofer Rind (Brookhaven National Laboratory)
- CHEP slides submitted for ATLAS committee review, thanks to all who provided input
- Special meeting held today to discuss ATLAS input for AF Data Delivery presentation at pre-CHEP workshop
- Doug and Oksana iterating on Analysis Portability slides
- Organized a BoF meeting in Norfolk on Saturday morning to follow up on CERNBox presentation at last week's HSF AF Forum
- Met on Monday to review/refocus ML analysis container development effort
- Shuwei gave a presentation to BNL NPPS this morning
- Critical OSG security alert affecting fuse mounts on RHEL8/9 (CVE)
-
14:15
Analysis Facilities - SLAC 5mSpeaker: Wei Yang (SLAC National Accelerator Laboratory (US))
-
14:20
Analysis Facilities - Chicago 5mSpeakers: Fengping Hu (University of Chicago (US)), Ilija Vukotic (University of Chicago (US))
-
14:10
-
14:25
→
14:35
AOB 10m
-
13:00
→
13:10