US ATLAS Computing Facility

US/Eastern
Description

Facilities Team Google Drive Folder

Zoom information

Meeting ID:  996 1094 4232

Meeting password: 125

Invite link:  https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09

 

 

    • 13:00 13:10
      WBS 2.3 Facility Management News 10m
      Speakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn Mc Kee (University of Michigan (US))
    • 13:10 13:20
      OSG-LHC 10m
      Speakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci

      Site admin office hours Tue Jan 25 1-4pm Central. Register here: https://docs.google.com/forms/d/e/1FAIpQLSdnvnv3uFdKN5MiVFmpFfsIYaZVZDLbpJUvTBprBsGpsSgKxQ/viewform

      Release

      • HTCondor-CE 5.1.3 (3.5 upcoming + 3.6)
      • CVMFS 2.9.0 (3.6 only)
      • CA cert updates (3.5 + 3.6)
      • HTCondor 9.5.0 (3.6 upcoming)
      • HTCondor 9.0.9 (3.5 upcoming + 3.6)

      Token Transition

      • osg-scitokens-mapfile-4-1 contains default ATLAS token -> local user mappings, please update!
      • If your site only supports ATLAS, GLOW, and OSG: you can update to OSG 3.6! https://opensciencegrid.org/docs/release/updating-to-osg-36/
      • CEs on token-supporting versions of HTCondor-CE (let me know if you expect to see your CE here but it isn't listed!)
        • bgk01.sdcc.bnl.gov
        • bgk02.sdcc.bnl.gov
        • gate01.aglt2.org
        • gate02.grid.umich.edu
        • gate04.aglt2.org
        • gpce03.fnal.gov
        • gpce04.fnal.gov
        • gridgk01.racf.bnl.gov
        • gridgk02.racf.bnl.gov
        • gridgk03.racf.bnl.gov
        • gridgk04.racf.bnl.gov
        • gridgk06.racf.bnl.gov
        • gridgk07.racf.bnl.gov
        • gridgk08.racf.bnl.gov
        • iut2-gk.mwt2.org
        • osg-gk.mwt2.org
        • spce01.sdcc.bnl.gov
        • spce02.sdcc.bnl.gov
        • uct2-gk.mwt2.org
      • CEs on old versions of HTCondor-CE
        • atlas-ce.bu.edu
        • gk01.atlas-swt2.org
        • gk04.swt2.uta.edu
        • grid1.oscer.ou.edu
        • mwt2-gk.campuscluster.illinois.edu
    • 13:20 13:50
      Topical Reports
      Convener: Robert William Gardner Jr (University of Chicago (US))
    • 13:50 13:55
      WBS 2.3.1 Tier1 Center 5m
      Speakers: Doug Benjamin (Brookhaven National Laboratory (US)), Eric Christian Lancon (Brookhaven National Laboratory (US))
    • 13:55 14:15
      WBS 2.3.2 Tier2 Centers

      Updates on US Tier-2 centers

      Convener: Fred Luehring (Indiana University (US))
      • It was a very good two weeks.
        • The generral draining on 1/13 was caused by two of the aipanda VMs becoming overloaded which resulted in HC offlining most grid sites. A user was uploading 500 MB tarball which caused the aipanda VMs to crash. The ADC team banned the user and restarted the affected VMs.
        • The draining of MWT2 today is for the quarterly scheduled preventative maintenance at the Illinois site.
      • I suspect at this point we won't need to further discussion of the site's readiness for run 3 because of the prior talk.
      • Get those purchases in!
      • 13:55
        AGLT2 5m
        Speakers: Philippe Laurens (Michigan State University (US)), Dr Shawn Mc Kee (University of Michigan (US)), Prof. Wenjing Wu (University of Michigan)

            Smooth running overall.

            1/5/2021
            One of the data switches in the UM Tier3 room (sw9-d-01) got stuck,
            and 6 work nodexs which are connected to the switch lost connection.
            The solution is to power cycle the switch.

            1/17/2022
            At 18:11, one dcache pool node (umfs06) rebooted by itself (not clear why).
            dcache was not restarted until 19 hours later, manually,
            which caused 500 jobs failure with stage-in errors.

            Currently finalizing quotes for end of cycle purchase.
            - R740xD2 with 24x 18TB disks
            - R6525 with AMD 7413

         

      • 14:00
        MWT2 5m
        Speakers: David Jordan (University of Chicago (US)), Jess Haney (Univ. Illinois at Urbana Champaign (US)), Judith Lorraine Stephen (University of Chicago (US))

        UC:

        • Trouble with two dcache nodes over the past couple weeks. Both are back up currently, but led to some job failures. Suspect that some hardware is bad on one. Will investigate when it's drained.
        • Second physical hardware move to new data center next week.
        • New compute and storage to arrive in the next few weeks.
        • Updated MWT2-TEST to osg 3.6

        IU

        • Updating compute nodes to condor 9
        • Machines moved to IU from UC are cabled and will start to get put back in production soon.

        UIUC

        • ICC PM today, site offline.
      • 14:05
        NET2 5m
        Speaker: Prof. Saul Youssef

         

        1. In the process of retiring 2.5 racks of 3TB storage (770 TB useable). 

        2. Added 4 nodes to GPFS xrootd cluster

        3. 1&2 greatly improved staging performance, GPFS slowness issues

        4. Remaining hardware orders finalizing through BU purchasing... 10 new transfer nodes and 3.8 PB NESE Ceph storage being purchased.   No new worker nodes.  

        5. 88 worker nodes schedule to arrive from DELL March 3

        6. Lining up collaborations and organization for bare metal cluster & UMass expansion. 

        7. Preparing to upgrade NET2-NESE networking to 400Gb/s.

        8. NESE Tape commissioning continues with NESE team, Xin, Alexei and ADC.

        Smooth operations in the past 2 weeks. 

      • 14:10
        SWT2 5m
        Speakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))

        SWT2_CPB -

        • Added second host to the webdav pool, plus upgraded XRootD to v5.4.0. Better performance and stability since implementing these changes.
        • Working with Hiro to optimize the concurrency settings in FTS.
        • Odd routing for transfers to RAL - not using LHCONE path?
        • Submitted quotes to our procurement office for the latest hardware purchase: 3 PB of storage, 4608 logical job slots in WN's.

        UTA_SWT2 -

        • Working with personnel at the data center to finalize our shutdown & hardware move dates.

        OU:

        - Not much to report, running stably.

        - Still occasional xrootd hangups, hopefully upgrading backend storage to 5.4.x will fix that.

        - Working with OU Purchasing to put Dell quote through, to spend remaining hardware funds.

         

    • 14:15 14:20
      WBS 2.3.3 HPC Operations 5m
      Speakers: Lincoln Bryant (University of Chicago (US)), Rui Wang (Argonne National Laboratory (US))

      TACC 

      • Filesystems issues over the past few days caused Harvester to crash. Seems resolved now. not sure if related to Ranch filesystem maintenance.
      • 72K SUs remaining

      NERSC

      • Maintenance day to transition to AY22 allocations
      • Focus is on Perlmutter now
    • 14:20 14:35
      WBS 2.3.4 Analysis Facilities
      Convener: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:20
        Analysis Facilities - BNL 5m
        Speaker: Ofer Rind (Brookhaven National Laboratory)
        • Rolling update to mount eos/user on Tier-3 interactive hosts this week
        • Preparing joint US AF presentation for S&C week
      • 14:25
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:30
        Analysis Facilities - Chicago 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
    • 14:35 14:55
      WBS 2.3.5 Continuous Operations
      Convener: Ofer Rind (Brookhaven National Laboratory)
      • Token readiness discussion at last week's WFMS weekly meeting
        • Follow up discussions at S&C week
      • Ongoing transfer issues at CPB (RAL IPV4 issue has been resolved but backlog still a problem)
      • Pilot 3 being deployed
      • Evaluating Run 3 readiness (see above)
      • 14:35
        US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5m
        Speakers: Mark Sosebee (University of Texas at Arlington (US)), Xin Zhao (Brookhaven National Laboratory (US))
      • 14:40
        Service Development & Deployment 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))

        Analytics

        • additional ES nodes drained for transport
        • three new logstash based ingresses
        • all the ML platform images are being upgraded today

        XCaches

        • fixes to gStream reports
        • testing ephemeral storage changes
        • some unexpected restarts due to k8s liveness probes failing.

        VP

        • running fine

        ServiceX

        • some developments got merged
        • needs more testing

         

      • 14:45
        Kubernetes R&D at UTA 5m
        Speaker: Armen Vartapetian (University of Texas at Arlington (US))

        Hardware setup for initial k8s cluster is not yet completed, as our admins were busy dealing with issues affecting the production cluster. Also discussions to finalize the shutdown & move of the UTA_SWT2 hardware.

    • 14:55 15:05
      AOB 10m