US ATLAS Computing Facility
Facilities Team Google Drive Folder
Zoom information
Meeting ID: 996 1094 4232
Meeting password: 125
Invite link: https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09
-
-
13:00
→
13:10
WBS 2.3 Facility Management News 10mSpeakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn Mc Kee (University of Michigan (US))
Happy New Year and welcome to 2023!
- We have a few milestones that are due and we need to have updates for them.
- Fred can summarize the procurement status and we need to finalize the "plan"
- ATLAS S&C is at the end of this month. Any items still to add in the demonstrators list?
- For this year we have work to do in planning and carrying out suitable mini-milestones for the March 2024 WLCG Data Challenge
-
13:10
→
13:20
OSG-LHC 10mSpeakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
- HTCondor 10 available in testing, seeking feedback!
- HTCondor 10.0.1 in osg-testing for EL7 and EL8
- HTCondor 10.2.0 in osg-upcoming-testing for EL7, EL8, and soon EL9
- Initial EL9 packages have been built and we're working through issues with the EL9 testing infrastructure
- Aiming for a release in February, certainly by the end of Q1
- Targeting packages in testing by late January or early February
- Latest apptainer RPM in EPEL 8 does not provide Singularity
- HTCondor 10 available in testing, seeking feedback!
- 13:20 → 13:25
-
13:25
→
13:45
WBS 2.3.2 Tier2 Centers
Updates on US Tier-2 centers
Convener: Fred Luehring (Indiana University (US))- Pretty good running over the holiday break with no major failures.
- Some ATLAS Monit plots have not been filling for the last few days but the sites appear to be running well.
- I had Mario Lassnig and Paul Nilsson make some change in the way transfers are reported to monit to solve a problem where 25%-50% of the transfers at a site were shown as unknown. They put the fix into production on Dec 8 and the unknown transfers disappeared from the monit transfer page but so did the transfers with protocols root and https. Therefore I do not completely trust today's transfer plots.
- I did finish the first draft of the global tier 2 procurement plan on Dec 19 and sent it to the Tier 2 PIs, I got no comments of any kind and there still are some issues that we have not reached a consensus on.
- The timing of the purchases Do we go for one big purchase in March with all FY22 and FY23 funds or do we split it into two purchases one in the next month and one near the end of the summer.
- What the target should be for the storage/compute split.
- A consistent cost estimation formula for the compute ($/kHS06) and ($/TB)
- Whether to go in on a single joint bid or to go separately. I guess separate but I's like to confirm.
- We should decide at next week's management meeting, if we are going to proceed with having each site make an operations plan.
- It seemed to me that there was not a complete consensus reached at the SLAC meeting.
- NET2
- LOCALGROUPDISK monitoring
-
13:25
AGLT2 5mSpeakers: Philippe Laurens (Michigan State University (US)), Dr Shawn Mc Kee (University of Michigan (US)), Prof. Wenjing Dronen
12/8 : new Kernel, FW, and Condor 9.0.17
New kernel ( 3.10.0-1160.80.1.el7) including a security issue.
We drained all our work nodes and interactive nodes in batches to reboot into this new kernel.
This was an opportunity to update the firmware (especially BIOS and Network Card that require reboot)
for the R630, C6420 and R6515 models of compute nodes. Condor was also updated from 9.0.16 to 9.0.17.
This whole process took over a week to complete.
During the draining period, BOINC jobs filled all empty job slots released by HTCondor.12/16 : gave up on VMware vSAN
After experimenting with the VSAN for a few months, we found out it is not reliable
for production usage on a 3 node vmware cluster, so we decided to delete VSAN from our vmware cluster12/18 : VMware 6 updates
Updated ESXi hosts to the latest vSphere 6.7.x and the ESXi hosts had all firmware updates applied.
Configured ESXi hosts to only advertise 1/2 of their AMD CPU cores
(to match the license requirements, aka not have to buy a second set of licenses).
Updated The TrueNAS systems to the Bluefin release.12/19 : One missing file was causing job failures, we declared file loss in rucio.
1/3/2023 : Vmware 7
updated to vmware 7 on the UM cluster.
MSU cluster update has been delayed by iSCSI network configuration, now planned for deployment this week.
Starting update to vmware 7 this week, some in parallel with iSCSI deployment, and concluding next week.
Ongoing:
Work with Dell for new quotes on R740xD2 storage, R6525 compute, one NVMe R7525 for MSU VMware
-
13:30
MWT2 5mSpeakers: David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Judith Lorraine Stephen (University of Chicago (US))
-
13:35
NET2 5mSpeakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)
NESE:
- A solution was found to create a flexible system with a dCache/xRootD part and a CEPH part.
- We would like to start installation of this additional part. Fred suggested start with the R740 machines currently not being used, so that we don't disturb the LOCALGROUPDISK at NESE
- Eduardo will reach out to Fred and Doug to make a plan
- A presentation is planned for the topical meeting in 2 week.
TRANSFER:
- On rack with 2021 machines has been disconnected and is being transferred.
- We would like to stop operations at NET2 so that BU can finish their part of the transfer.
- A total right now of 5 racks are being transferred, but only 2 can be connected at this point.
NEW RACKS:
- No updates at this point.
-
13:40
SWT2 5mSpeakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))
UTA
- Quiet operations over the holiday break
- Work is progressing on the network replacement; defining the configurations.
- Hope to take a downtime in the next few weeks to physically replace the switches.
- The switch replacements will allow us to merge the K8s cluster into the main system
OU
Working well over the break.
Had brief xrootd glitch when one data server disappeared from the network.
Fixed by moving that to a different network switch port.
- Pretty good running over the holiday break with no major failures.
-
13:45
→
13:50
WBS 2.3.3 HPC Operations 5mSpeakers: Lincoln Bryant (University of Chicago (US)), Rui Wang (Argonne National Laboratory (US))
- NERSC
- Cori running fine, almost done with allocation. 0.3% remaining.
- Perlmutter online, will use "overrun" QOS until the end of the period. Have it set to "mintime=3600" as per Rod's suggestion to pick up some whole-node 10k sim jobs. Jedi hasn't placed any in the queue, though.
- TACC
- Successfully running jobs at a small scale (1 node) with CVMFSExec. "ONLINE" in PanDA, 0 job failures over night.
- General
- Proxies not autorenewing at NERSC or TACC. Have to restart harvester w/ new proxy daily. I think we need a new long-lived proxy in PanDA. Re-uploaded long-lived proxy to CERN MyProxy and made Jira ticket.
- NERSC
-
13:50
→
14:05
WBS 2.3.4 Analysis FacilitiesConveners: Ofer Rind (Brookhaven National Laboratory), Wei Yang (SLAC National Accelerator Laboratory (US))
-
13:50
Analysis Facilities - BNL 5mSpeaker: Ofer Rind (Brookhaven National Laboratory)
- Interesting presentations at IRIS-HEP AGC Demo Day
- Working on DASK integration and container build infrastructure
-
13:55
Analysis Facilities - SLAC 5mSpeaker: Wei Yang (SLAC National Accelerator Laboratory (US))
- 14:00
-
13:50
-
14:05
→
14:25
WBS 2.3.5 Continuous OperationsConvener: Ofer Rind (Brookhaven National Laboratory)
-
14:05
US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5mSpeaker: Mark Sosebee (University of Texas at Arlington (US))
-
14:10
Service Development & Deployment 5mSpeakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
XCache
- running fine
- BHAM and OXFORD had suboptimal settings for Block size and prefetch. Asked them to fix.
- will have a new xcache at PerlMutter today or tomorrow
- preliminary analysis showed that pmerge jobs basically never reuse data. So these datasets will be not given virtual placement.
VP
- running fine
- BNL VP now rampped from 95 to 500 cores.
- over the holidays I updated everything: node.js, redis, redis client, made everything run asynch.
- now not creating VP placements for panda.um.* datasets
Varnish
- AGLT2 instance running fine.
- will create 4 new instances at MWT2 UC and IU (2 for Frontier and 2 for CVMFS) and make them first choice
- Will work with the SLATE team to fine tune the Helm chart, pull image from Harbor, and provide some independent testing and documentation.
-
14:15
Kubernetes R&D at UTA 5mSpeaker: Armen Vartapetian (University of Texas at Arlington (US))
- Trying to understand how optimized is the job CPU requests coefficient sent from Harvester (has 0.9 scale down value). The idea of it is to leave CPU request space for other system/auxiliary pods. The issue was, that due to that, the K8S scheduler often was managing to squeeze in several more SCORE jobs on top of the available core count (basically overcommitting the node). I pinged Fernando about this, and after checking, he noticed he has the same issue in his Google cloud as well. So, this needs to be a bit optimized.
- During December there were a bunch of tasks with SCORE_HIMEM jobs, which were pushing out the MCORE jobs, stuck in activated state. That was quite strange as we had limit on number of running SCORE_HIMEM jobs (similar to limit on running SCORE jobs) to avoid such behavior. I noticed that Victoria is also suffering from the same issue (running only SCORE_HIMEM, all MCORE jobs stuck). Noticed the issue is the name of the parameter "resource_type_limits.SCORE_HIMEM" in CRIC, which got an extra space typo in the name (no idea how it got there). After the fix, things went back to normal.
- SWT2_CPB_K8S started to drain on New Year's eve, and Jan.1 we had a lot of activated jobs, but nothing running. On K8S side all was fine and healthy. Looking in the Harvester, I saw the workers running, but the logs show it was failed to get pilot code. One possible thing to try (as there was no expert help available) to remove the pilot url path in CRIC, which we were using for a pre-release pilot version with a fix (now in production), and that did the trick, things started to run again.
-
14:05
-
14:25
→
14:35
AOB 10m
-
13:00
→
13:10