US ATLAS Computing Facility

US/Eastern
Description

Facilities Team Google Drive Folder

Zoom information

Meeting ID:  996 1094 4232

Meeting password: 125

Invite link:  https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09

 

 

    • 13:00 13:10
      WBS 2.3 Facility Management News 10m
      Speakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn Mc Kee (University of Michigan (US))

      Happy New Year and welcome to 2023!

       

       

       

    • 13:10 13:20
      OSG-LHC 10m
      Speakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
      • HTCondor 10 available in testing, seeking feedback!
        • HTCondor 10.0.1 in osg-testing for EL7 and EL8
        • HTCondor 10.2.0 in osg-upcoming-testing for EL7, EL8, and soon EL9
      • Initial EL9 packages have been built and we're working through issues with the EL9 testing infrastructure
        • Aiming for a release in February, certainly by the end of Q1
        • Targeting packages in testing by late January or early February
      • Latest apptainer RPM in EPEL 8 does not provide Singularity
    • 13:20 13:25
      WBS 2.3.1 Tier1 Center 5m
      Speakers: Doug Benjamin (Brookhaven National Laboratory (US)), Eric Christian Lancon (Brookhaven National Laboratory (US))
      • Allocating more jobs slots to VP queue (up to 500)
      •  
    • 13:25 13:45
      WBS 2.3.2 Tier2 Centers

      Updates on US Tier-2 centers

      Convener: Fred Luehring (Indiana University (US))
      • Pretty good running over the holiday break with no major failures.
        • Some ATLAS Monit plots have not been filling for the last few days but the sites appear to be running well.
        • I had Mario Lassnig and Paul Nilsson make some change in the way transfers are reported to monit to solve a problem where 25%-50% of the transfers at a site were shown as unknown. They put the fix into production on Dec 8 and the unknown transfers disappeared from the monit transfer page but so did the transfers with protocols root and https. Therefore I do not completely trust today's transfer plots.
      • I did finish the first draft of the global tier 2 procurement plan on Dec 19 and sent it to the Tier 2 PIs, I got no comments of any kind and there still are some issues that we have not reached a consensus on.
        • The timing of the purchases Do we go for one big purchase in March with all FY22 and FY23 funds or do we split it into two purchases one in the  next month and one near the end of the summer.
        • What the target should be for the storage/compute split.
        • A consistent cost estimation formula for the compute ($/kHS06) and ($/TB)
        • Whether to go in on a single joint bid or to go separately. I guess separate but I's like to confirm.
      • We should decide at next week's management meeting, if we are going to proceed with having each site make an operations plan.
        • It seemed to me that there was not a complete consensus reached at the SLAC meeting.
      • NET2
      • LOCALGROUPDISK monitoring
      • 13:25
        AGLT2 5m
        Speakers: Philippe Laurens (Michigan State University (US)), Dr Shawn Mc Kee (University of Michigan (US)), Prof. Wenjing Dronen

        12/8 : new Kernel, FW, and Condor 9.0.17
         New kernel ( 3.10.0-1160.80.1.el7) including a security issue.
         We drained all our work nodes and interactive nodes in batches to reboot into this new kernel.
         This was an opportunity to update the firmware (especially BIOS and Network Card that require reboot)
         for the R630, C6420 and R6515 models of compute nodes. Condor was also updated from 9.0.16 to 9.0.17.
         This whole process took over a week to complete.
         During the draining period, BOINC jobs filled all empty job slots released by HTCondor.

         

        12/16 : gave up on VMware vSAN
         After experimenting with the VSAN for a few months, we found out it is not reliable
         for production usage on a 3 node vmware cluster, so we decided to delete VSAN from our vmware cluster

         

        12/18 : VMware 6 updates
         Updated ESXi hosts to the latest vSphere 6.7.x and the ESXi hosts had all firmware updates applied.
         Configured ESXi hosts to only advertise 1/2 of their AMD CPU cores
         (to match the license requirements, aka not have to buy a second set of licenses).
         Updated The TrueNAS systems to the Bluefin release.

         

        12/19 : One missing file was causing job failures, we declared file loss in rucio.

         

        1/3/2023 : Vmware 7
         updated to vmware 7 on the UM cluster.
         MSU cluster update has been delayed by iSCSI network configuration, now planned for deployment this week.
         Starting update to vmware 7 this week, some in parallel with iSCSI deployment, and concluding next week.
         

        Ongoing:
         Work with Dell for new quotes on R740xD2 storage, R6525 compute, one NVMe R7525 for MSU VMware
         

      • 13:30
        MWT2 5m
        Speakers: David Jordan (University of Chicago (US)), Farnaz Golnaraghi (University of Chicago (US)), Judith Lorraine Stephen (University of Chicago (US))

        Mostly quiet over the break.

        UC squid went down for a few days due to filesystem issues.

        Discussing procurement and upgrade schedule for the next couple of months.

      • 13:35
        NET2 5m
        Speakers: Eduardo Bach (University of Massachusetts (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US)), William Axel Leight (University of Massachusetts Amherst)

        NESE:

         

        • A solution was found to create a flexible system with a dCache/xRootD part and a CEPH part.
        • We would like to start installation of this additional part. Fred suggested start with the R740 machines currently not being used, so that we don't disturb the LOCALGROUPDISK at NESE
        • Eduardo will reach out to Fred and Doug to make a plan
        • A presentation is planned for the topical meeting in 2 week.

         

        TRANSFER:

         

        • On rack with 2021 machines has been disconnected and is being transferred.
        • We would like to stop operations at NET2 so that BU can finish their part of the transfer.
        • A total right now of 5 racks are being transferred, but only 2 can be connected at this point.

         

        NEW RACKS:

         

        • No updates at this point.
      • 13:40
        SWT2 5m
        Speakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))

        UTA

        • Quiet operations over the holiday break
        • Work is progressing on the network replacement; defining the configurations.  
          • Hope to take a downtime in the next few weeks to physically replace the switches.
          • The switch replacements will allow us to merge the K8s cluster into the main system

        OU

        Working well over the break.

        Had brief xrootd glitch when one data server disappeared from the network.

        Fixed by moving that to a different network switch port.

         

    • 13:45 13:50
      WBS 2.3.3 HPC Operations 5m
      Speakers: Lincoln Bryant (University of Chicago (US)), Rui Wang (Argonne National Laboratory (US))

       

      • NERSC
        • Cori running fine, almost done with allocation. 0.3% remaining.
        • Perlmutter online, will use "overrun" QOS until the end of the period. Have it set to "mintime=3600" as per Rod's suggestion to pick up some whole-node 10k sim jobs. Jedi hasn't placed any in the queue, though. 
      • TACC
        • Successfully running jobs at a small scale (1 node) with CVMFSExec. "ONLINE" in PanDA, 0 job failures over night.
      • General
        • Proxies not autorenewing at NERSC or TACC. Have to restart harvester w/ new proxy daily. I think we need a new long-lived proxy in PanDA. Re-uploaded long-lived proxy to CERN MyProxy and made Jira ticket. 
    • 13:50 14:05
      WBS 2.3.4 Analysis Facilities
      Conveners: Ofer Rind (Brookhaven National Laboratory), Wei Yang (SLAC National Accelerator Laboratory (US))
      • 13:50
        Analysis Facilities - BNL 5m
        Speaker: Ofer Rind (Brookhaven National Laboratory)
        • Interesting presentations at IRIS-HEP AGC Demo Day
        • Working on DASK integration and container build infrastructure
      • 13:55
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:00
        Analysis Facilities - Chicago 5m
        Speakers: Fengping Hu (University of Chicago (US)), Ilija Vukotic (University of Chicago (US))
        • Added a harbor proxy cache service
        • Added the jwt token to the coffea-casa notebook pod
    • 14:05 14:25
      WBS 2.3.5 Continuous Operations
      Convener: Ofer Rind (Brookhaven National Laboratory)
      • BNL_ATLAS_VP increased to 500 slots
      • OU planning update to OSG 3.6 EL9 testing release of HTCondor-CE by the end of this month
      • Begin NET2 decommissioning process - CRIC updates, data migration
      • 14:05
        US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5m
        Speaker: Mark Sosebee (University of Texas at Arlington (US))
      • 14:10
        Service Development & Deployment 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))

        XCache

        • running fine
        • BHAM and OXFORD had suboptimal settings for Block size and prefetch. Asked them to fix.
        • will have a new xcache at PerlMutter today or tomorrow
        • preliminary analysis showed that pmerge jobs basically never reuse data. So these datasets will be not given virtual placement.

        VP

        • running fine
        • BNL VP now rampped from 95 to 500 cores.
        • over the holidays I updated everything: node.js, redis, redis client, made everything run asynch.
        • now not creating VP placements for panda.um.* datasets 

        Varnish

        • AGLT2 instance running fine.
        • will create 4 new instances at MWT2 UC and IU (2 for Frontier and 2 for CVMFS) and make them first choice
        • Will work with the SLATE team to fine tune the Helm chart, pull image from Harbor, and provide some independent testing and documentation.

         

      • 14:15
        Kubernetes R&D at UTA 5m
        Speaker: Armen Vartapetian (University of Texas at Arlington (US))
        • Trying to understand how optimized is the job CPU requests coefficient sent from Harvester (has 0.9 scale down value). The idea of it is to leave CPU request space for other system/auxiliary pods. The issue was, that due to that, the K8S scheduler often was managing to squeeze in several more SCORE jobs on top of the available core count (basically overcommitting the node). I pinged Fernando about this, and after checking, he noticed he has the same issue in his Google cloud as well. So, this needs to be a bit optimized.
        • During December there were a bunch of tasks with SCORE_HIMEM jobs, which were pushing out the MCORE jobs, stuck in activated state. That was quite strange as we had limit on number of running SCORE_HIMEM jobs (similar to limit on running SCORE jobs) to avoid such behavior. I noticed that Victoria is also suffering from the same issue (running only SCORE_HIMEM, all MCORE jobs stuck). Noticed the issue is the name of the parameter "resource_type_limits.SCORE_HIMEM" in CRIC, which got an extra space typo in the name (no idea how it got there). After the fix, things went back to normal.
        • SWT2_CPB_K8S started to drain on New Year's eve, and Jan.1 we had a lot of activated jobs, but nothing running. On K8S side all was fine and healthy. Looking in the Harvester, I saw the workers running, but the logs show it was failed to get pilot code. One possible thing to try (as there was no expert help available) to remove the pilot url path in CRIC, which we were using for a pre-release pilot version with a fix (now in production), and that did the trick, things started to run again.
    • 14:25 14:35
      AOB 10m