US ATLAS Computing Facility

US/Eastern
Description

Facilities Team Google Drive Folder

Zoom information

Meeting ID:  996 1094 4232

Meeting password: 125

Invite link:  https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09

 

 

    • 13:00 13:10
      WBS 2.3 Facility Management News 10m
      Speakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn Mc Kee (University of Michigan (US))
    • 13:10 13:20
      OSG-LHC 10m
      Speakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
    • 13:20 13:50
      Topical Reports
      Convener: Robert William Gardner Jr (University of Chicago (US))
    • 13:50 13:55
      WBS 2.3.1 Tier1 Center 5m
      Speakers: Doug Benjamin (Brookhaven National Laboratory (US)), Eric Christian Lancon (Brookhaven National Laboratory (US))

      business as usual

    • 13:55 14:15
      WBS 2.3.2 Tier2 Centers

      Updates on US Tier-2 centers

      Convener: Fred Luehring (Indiana University (US))
      • Sites generally running well.
        • NET2.0 (Legacy site) running with about 2/3.
      • NET2.1 prgressing/
      • 13:55
        AGLT2 5m
        Speakers: Philippe Laurens (Michigan State University (US)), Dr Shawn Mc Kee (University of Michigan (US)), Prof. Wenjing Dronen

        Smooth running.

        UM site is draining the work nodes and dcache nodes to apply firmware updates in batches
        MSU site will follow

      • 14:00
        MWT2 5m
        Speakers: David Jordan (University of Chicago (US)), Judith Lorraine Stephen (University of Chicago (US))

        IU network maintenance October 5/6. New switch gear is now in production

        Ran into a configuration issue with our condor group quotas getting ignored by condor. To be debugged further.

        Still planning to upgrade condor on our central managers.

      • 14:05
        NET2 5m
        Speakers: Rafael Antonio Lopez, William Axel Leight (University of Massachusetts Amherst)
      • 14:10
        SWT2 5m
        Speakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))

        UTA:

        • Problem with SLURM database last week knocked us offline.   DB tables grew too large for their partition and prevented jobs being accepted.  Table usage was reduced with dump/restore and moved to a new partition.  Should not recur.
        • Problem with a compute node where CVMFS failed in an odd manner.  Caused a GGUS ticket.  Internal checks should have removed the node from production, but failed for unknown reasons; investigating.
        • Working to meet with OIT Networking to acknowledge the LHCOne AUP,   Having issues agreeing on time.  Netsite was created for SWT2_CPB and replicates the information that was under UTA_SWT2.
        • IPV6 work on the CE's is being tested.  Need SLATE infrastructure to be updated to work.

        OU:

        • All running well.
        • Had a brief glitch last week, where electrical work took down some compute nodes, which caused lost heartbeat failures.

         

    • 14:15 14:20
      WBS 2.3.3 HPC Operations 5m
      Speakers: Lincoln Bryant (University of Chicago (US)), Rui Wang (Argonne National Laboratory (US))
    • 14:20 14:35
      WBS 2.3.4 Analysis Facilities
      Conveners: Ofer Rind (Brookhaven National Laboratory), Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:20
        Analysis Facilities - BNL 5m
        Speaker: Ofer Rind (Brookhaven National Laboratory)
        • IRIS-HEP AGC planning meeting last week, discussed OKD support and site support/testing
      • 14:25
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:30
        Analysis Facilities - Chicago 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
    • 14:35 14:55
      WBS 2.3.5 Continuous Operations
      Convener: Ofer Rind (Brookhaven National Laboratory)
      • BNL OSG GK's updated to HTCondor-CE 5.1.5; upgrade of ATLAS GK's in progress
      • File transfer timeouts in ANALY_BNL_VP - still investigating 
      • WT2 deletion timeouts (ggus)
      • Status of SWT2 NetSite update in CRIC? (ggus)
      • Pilot 3.4.0.118 released, includes PANDA token support and a lot more:  https://github.com/PanDAWMS/pilot3/releases/tag/3.4.0.118
      • GDB this morning, including summary of the pre-GDB Authz and IAM workshop
      • 14:35
        US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5m
        Speakers: Mark Sosebee (University of Texas at Arlington (US)), Xin Zhao (Brookhaven National Laboratory (US))
      • 14:40
        Service Development & Deployment 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
      • 14:45
        Kubernetes R&D at UTA 5m
        Speaker: Armen Vartapetian (University of Texas at Arlington (US))
        • The K8S cluster was stable past weeks running production jobs.
        • Noticed some inefficiency in completely filling the cores in older machines. Appeared to be issue with pod scheduling, when overall available node memory was slightly lower (~0.3%) than the requested memory (2000MiB/16000MiB for SCORE/MCORE jobs).
        • Discussion with Fernando about possible solution. The easiest fix was to slightly scale down the memory request. Fernando introduced a new parameter in the Harvester: k8s.resources.requests.memory_scheduling_ratio. It was set to 98% for SWT2_CPB_K8S site in CRIC.
        • That was introduced Friday (Oct.9), and after that the cluster core occupancy level improved from around 500 cores to close to 600 cores (see the attached plot).
        • Earlier we put in CRIC limit on SCORE jobs (otherwise SCORE jobs gradually occupied most of the cores). We have reasonable mixture of SCORE/MCORE jobs now. To understand how to optimize the limits and maybe the method itself.
        • Also understand K8s methods which possibly can be used to optimize core occupancy for various SCORE/MCORE job mixtures.

         

    • 14:55 15:05
      AOB 10m