US ATLAS Computing Facility

US/Eastern
Description

Facilities Team Google Drive Folder

Zoom information

Meeting ID:  996 1094 4232

Meeting password: 125

Invite link:  https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09

 

 

    • 13:00 13:10
      WBS 2.3 Facility Management News 10m
      Speakers: Robert William Gardner Jr (University of Chicago (US)), Dr Shawn McKee (University of Michigan ATLAS Group)
    • 13:10 13:20
      OSG-LHC 10m
      Speakers: Brian Lin (University of Wisconsin), Matyas Selmeci

      Packages Ready for Testing

      • Frontier Squid 4.13-5.2 (also available in the opensciencegrid/frontier-squid:testing container!)
      • HTCondor 9.0.0, BLAHP 2.0.1 (available in OSG 3.5 upcoming testing)
      • HTCondor-CE 5.1.0 (available in OSG 3.5 upcoming testing)
      • VOMS 2.0.16: add support for signing VOMS proxies through an IAM endpoint. Is the ATLAS IAM endpoint (https://atlas-auth.web.cern.ch/?) ready for testing?

      Misc

      • Any sites exploring EL8?
      • HTCondor Week highly recommended for HTCondor sites (May 24-28)! https://agenda.hep.wisc.edu/event/1579/
      • Does US ATLAS need to split pilots between multiple users?
        "/atlas/*Role=production/Capability=*" usatlas1
        "/atlas/*Role=lcgadmin/Capability=*" usatlas2
        "/atlas/*Role=software/Capability=*" usatlas2
        "/atlas/*" usatlas3
    • 13:20 13:35
      Topical Reports
      Convener: Robert William Gardner Jr (University of Chicago (US))
      • 13:20
        TBD 15m
    • 13:35 13:40
      WBS 2.3.1 Tier1 Center 5m
      Speakers: Eric Christian Lancon (CEA/IRFU,Centre d'etude de Saclay Gif-sur-Yvette (FR)), Doug Benjamin (Duke University (US))
    • 13:40 14:00
      WBS 2.3.2 Tier2 Centers

      Updates on US Tier-2 centers

      Convener: Fred Luehring (Indiana University (US))
      • Not so smooth running:


        • MWT2 had a couple of maintenance periods (one at UIUC and one at IU) followed by a cooling incident at UC.
        • Over the weekend and into Monday, aipanda158 misbehaved and partially drained MWT2.
        • Bad user jobset at AGLT2, caused ~33k failures over the weekend. Did not block user correctly. Scout jobs missed the issue but the failing jobs were retried up to 15 times.
      • Slowly making progress on IPV6. IU site close to done. SWT2_UTA close.
      • Still struggling with setting downtimes.
      • Don't know status of work in XRootD to enable TPC.
      • 13:40
        AGLT2 5m
        Speakers: Philippe Laurens (Michigan State University (US)), Dr Shawn McKee (University of Michigan ATLAS Group), Prof. Wenjing Wu (Computer Center, IHEP, CAS)

            Ticket 151358 solved.
            About file transfer errors.  Root cause was fiber cut. Repaired after a couple days.
            The UM router lost its default route.
            Failover to backup Merit path worked.
            IPv4 connecttivity was not affected.  
            We were missing some IPv6 route announcements to non-LHCONE sites.
            The campus networking upgrades at UM and MSU will provide better control.

            Problem with one user jobs: fraction of job with failures and many retries, not detected by scouts.
            Unfortunate timing on a Saturday.
            Happened to be an MSU student, but a coincidence.
            Tried to temporarily ban user, but did not work.
            Jobs eventually stopped.
            User added code to detect and avoid problem in future.  
            But underlying systemic weakness not resolved.

            Resolved a BOINC issue. Only seen on most UM WNs.
            A number of rootfs-* folders were appearing in /tmp.
            Those were created when uncompressing container images.
            Causing issue after we had moved the boinc workspace out of /tmp into its own partition.
            /tmp now smaller and was getting full.
            Noticed correlation of these files being present on nodes without singularity locally installed.
            While BOINC usese singularity from /cvmfs.  
            So the exact mechanism is not clear, but installing singularity resolved issue.

         

      • 13:45
        MWT2 5m
        Speakers: David Jordan (University of Chicago (US)), Judith Lorraine Stephen (University of Chicago (US))

        Cooling failure in the UChicago server room took storage and UC compute offline yesterday/today. Working on bringing services back online

        HTTP-TPC GGUS ticket finally closed

        Kernels updated on IU compute, seems to have helped timeout issues at IU that have been taking us offline by hammercloud

        IU network maintenance ongoing

        UIUC PM last Wednesday. All UIUC compute nodes are now running licensed RedHat 7

      • 13:50
        NET2 5m
        Speaker: Prof. Saul Youssef (Boston University (US))

         

        Smooth operations in the past two weeks.

        Some 10 year old CPUs failing.  Planning to buy more soon.

        Top priority:  Testing xrd..   Containerized xrd 5.1.1 works for GPFS export.  Some problems reported back to Wei and Andy, will test new releases as soon as available.  

      • 13:55
        SWT2 5m
        Speakers: Dr Horst Severini (University of Oklahoma (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))

        OU:

        -Nothing to report, all running well

        UTA:

        Now running XRootD HTTP-TPC test instance from Wei.

        Smooth Production

    • 14:00 14:05
      WBS 2.3.3 HPC Operations 5m
      Speakers: Doug Benjamin (Duke University (US)), lincoln bryant

      TACC going along slowly

      NERSC had monthly downtime. - ran out of work (assigned tasks) when reduced the throughput.

      Lincoln still sorting out credential renewal with Tadashi. When his credentials are automatically renewed and test jobs run successfully, then NERSC becomes his problem

    • 14:05 14:20
      WBS 2.3.4 Analysis Facilities
      Convener: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:05
        Analysis Facilities - BNL 5m
        Speaker: William Strecker-Kellogg (Brookhaven National Lab)

        NTR

      • 14:10
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))

        NTR

      • 14:15
        Analysis Facilities - Chicago 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))

        All running fine. Will need to add ephemeral storage request, some private jupyter labs are unexpectedly running out of it.

    • 14:20 14:40
      WBS 2.3.5 Continuous Operations
      Convener: Ofer Rind
      • Still myriad issues/confusion with process of declaring downtimes (SWT2, BU, MWT2,....)
        • OSG topology: severity, protocol/service matching
        • CRIC configuration (e.g. BU SE's)
        • HC/Switcher roles, delays
      • How to ban a problematic user?  Contact DPA
      • 14:20
        US Cloud Operations Summary: Site Issues, Tickets & ADC Ops News 5m
        Speakers: Mark Sosebee (University of Texas at Arlington (US)), Xin Zhao (Brookhaven National Laboratory (US))
      • 14:25
        Service Development & Deployment 5m
        Speakers: Ilija Vukotic (University of Chicago (US)), Robert William Gardner Jr (University of Chicago (US))
        • BNL XCache (v5.1.1) now sending heartbeats; need to investigate issue with VP queue
        • Set up of XRootd HTTP-TPC testbeds at SWT2 and BNL is progressing

         

        XCaches

        • all working fine.
        • work on understanding gStream packet loss
        • trying to understand xcache accounting of where the data comes from.
        • will be talking to Beijing site to get SLATE there, squid, and two xcaches.

        VP

        • all queues working fine.
        • jobs efficiency is high

        Squids

        • No meeting today
        • All squids were working fine. 
        • Improvements in squid config seems to be working fine at UC, should be rolled wider by next week. It would require getting one of the disks out from xcache and giving it to squid.

        Alarm & Alert service

        • new way to subscribe to alert emails
        • still some development to be done.
    • 14:40 14:45
      AOB 5m