US ATLAS Computing Facility

US/Eastern
Description

Facilities Team Google Drive Folder

Zoom information

Meeting ID:  996 1094 4232

Meeting password: 125

Invite link:  https://uchicago.zoom.us/j/99610944232?pwd=ZG1BMG1FcUtvR2c2UnRRU3l3bkRhQT09

 

 

    • 13:00 13:05
      WBS 2.3 Facility Management News 5m
      Speakers: Alexei Klimentov (Brookhaven National Laboratory (US)), Dr Shawn Mc Kee (University of Michigan (US))

      Working on discussion of USATLAS Ops https://docs.google.com/document/d/1pjwG1LAjOPWsSrdas4WvYYfoyn8NyLuk5u4L6opNb4Q/edit#heading=h.o4xh9deafqh6

      Trying to finalize the HTC24 agenda, see https://docs.google.com/document/d/1em3rfH8lSwa5HnUomXwLKsSXFQGAV6S3pvT94ld-RY4/edit#heading=h.v2jvxcc92986

      We plan to have a pre-scrubbing meeting on or shortly after June 20, 2024 with a goal of getting solid draft slides for the July scrubbing.

       

    • 13:05 13:10
      OSG-LHC 5m
      Speakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci

      Software

    • 13:10 13:30
      WBS 2.3.1: Tier1 Center
      Convener: Alexei Klimentov (Brookhaven National Laboratory (US))
      • 13:10
        Tier-1 Infrastructure 5m
        Speaker: Jason Smith

        OKD cluster is available for ATLAS and finishing evaluation of OpenShift.  OKD will be retired and hardware will be reused for OpenShift.

      • 13:15
        Compute Farm 5m
        Speaker: Thomas Smith

        Lots of work behind the scenes on a vertical slice, with two managers and five worker nodes, all AlmaLinux9, dual-stacked.

      • 13:20
        Storage 5m
        Speakers: Carlos Fernando Gamboa (Department of Physics-Brookhaven National Laboratory (BNL)-Unkno), Carlos Fernando Gamboa (Brookhaven National Laboratory (US))
        • Secure xroot support enabled at doors in dCache 9.2.17, including bug fix [www.dcache.org #10562]:Pools do not reload updated certificates.

        • Awaiting JBODs for new pool nodes (ETA 05/28) 14PB total to be deployed; Head Nodes delivered.

        • Rolling OS upgrade to RHEL 8 for warranted pool servers.

          • Using 2PB buffer for transparent data migration and OS upgrades; 15PB left.

        • Tuning DMZ pool activities ongoing.

          • MCTAPE write activity increased.

      • 13:25
        Tier1 Operations and Monitoring 5m
        Speaker: Ivan Glushkov (University of Texas at Arlington (US))

         

        • Added two EL9/Condor 21 CEs to OSG topology and CRIC
        • Blacklisted (05/10) due to running out of file descriptors. This should be solved by EL9/cvmfs client update in the near future.
        • Hi squid usage. “This is the new norm”
        • BNL Tape firmware update: Allowed us to define and test blacklisting chain OSG/CRC for TAPE REST API.
    • 13:30 13:40
      WBS 2.3.2 Tier2 Centers

      Updates on US Tier-2 centers

      Conveners: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
      • Reasonable running over the last 4 weeks when there was sufficient work.
        • CERN did run out of work on two occasions causing significant draining
        • Frontier error messages related to Varnish caused successful jobs to be marked a failed at AGLT2 & MWT2.
        • NET2 ran pretty good with some minor draining incidents.
          • Scheduled downtime today for power work..
          • 3 more PB came online and is slowly filling.
        • CPB managed to remove LSM but there was on knock-on effects.
          • Ran into an XRootD issue caused by upper/lower case in web addresses that made transfers to sites running storm.
          • They are working with Wei on getting the storage tokens enabled.
        • Taiwan Tier 1 (TW-FTT) was supposed to be put online this week when network errors seemed to have been solved.
          • Did not seem to happen. The site is still shown in test.
      • Judith held a training session for Foreman / Puppet training for Aidan Rosberg (IU) and Zach Booth (CPB).
      • Long discussion of coordinating cvmfs debugging for various problems seen at various sites.
        • Kaushik was worried that the check that cvmfs is ok in the pilot wrapper before starting the pilot may cause issues.
    • 13:40 13:50
      WBS 2.3.3 Heterogenous Integration and Operations (HIOPS)

      HIOPS

      Convener: Rui Wang (Argonne National Laboratory (US))
      • 13:40
        HPC Operations 5m
        Speaker: Rui Wang (Argonne National Laboratory (US))

        TACC

        • HC jobs are running smoothly, brought the queue online today
        • The queuing time is very long (2 hour job ~ few hours; 6 hour job ~1days; 10 hour job ~3-4days)
          • Trying the job packing of 2 jobs x 3 nodes (6 hour) for worker
        • upgrade atlas-cvmfsexec to 1.0.27 with cvmfsexec v4.39

        Perlmutter

        • 118K CPU hours added (total 818K)
      • 13:45
        Integration of Complex Workflows on Heterogeneous Resources 5m
        Speaker: Doug Benjamin (Brookhaven National Laboratory (US))
    • 13:50 14:10
      WBS 2.3.4 Analysis Facilities
      Conveners: Ofer Rind (Brookhaven National Laboratory), Wei Yang (SLAC National Accelerator Laboratory (US))
      • 13:50
        Analysis Facilities - BNL 5m
        Speaker: Dr Quilan Huang (BNL)
      • 13:55
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:00
        Analysis Facilities - Chicago 5m
        Speaker: Fengping Hu (University of Chicago (US))
        • starting development on OpenAI assistant for AF.
          • need to collect documentation and discourse conversations
        • adding GPU info  and condor queue info to AF monitoring
        • Highlights of AF maintenence done on Monday/Tuesday
          • Migration of all servers to Enterprise Linux 9
          • HTCondor updated to the OSG 23 release series
          • Update of the /data filesystem to Ceph v17
          • System BIOS and firmware updates on all servers
    • 14:10 14:25
      WBS 2.3.5 Continuous Operations
      Convener: Ofer Rind (Brookhaven National Laboratory)
      • DC24 Report will be submitted next week, so please consider giving it a once-over if you haven't already (link)
      • 14:10
        ADC Operations, US Cloud Operations: Site Issues, Tickets & ADC Ops News 5m
        Speaker: Ivan Glushkov (University of Texas at Arlington (US))
        • ADC:
          • ~A week of low GRID occupancy due to lack of simulation in the system
        • Rucio:
          • xrootd to storm sites - transfers fail due to xrootd bug. It will be fixed in version 5.7 this summer. Temporary patch is removing distances between these sites.
        • HC:
          • Short mass-blaklisting earlier today due to a Panda bug.
        • CVMFS Monitoring - available, but consists of two separate categories of errors - not avle to get pilot and not being able to access the cvmfs
        • Started Meetings:
          • First ADC Fabrics Meeting (Indico:1414901)
            • “to facilitate communication and contributions between ADC and infrastructure providers”. Monthly
          • US ATLAS Distributed Computing Ops (Doc)
            • Same idea as the ADC Ops meeting.
            • Daily, 9:30 AM CDT / 10:30 AM EDT / 4:30 PM CEST / 10:30 PM CST
            • Started this Monday with trail period of one week.
            • Already addressed: SWT2 Storage Tokens, MWT2 CA related transfer errors, monitoring for CVMFS issues, etc.
      • 14:15
        Services DevOps 5m
        Speaker: Ilija Vukotic (University of Chicago (US))
        • XCache & VP
          • MWT2 xcaches back in operation. Reactivated VP queue today
          • Still having 8 xcaches serving AF
          • VP working fine
        • Varnishes
          • added an instance in NRP for NET2. Works fine.
          • will try configuring DNS Anycast in Cloudflare so all the varnishes would be behind the same name. Failovers would be automatic. 5$/node/month.  
        • ServiceX
          • Production instance on AF restarted. Still running with a lot of manual fixes.
          • Today moving all the testing instances to River-dev cluster
      • 14:20
        Facility R&D 5m
        Speaker: Lincoln Bryant (University of Chicago (US))
        • Successful Kubernetes tutorial and workshop last month, a lot of interesting ground covered
          • 20+ successful single-node K8S clusters built with Kubespray
          • Kueue (multi-user fair share scheduling), Karmada (multi-cluster), stretched K8S over Wireguard VPN, Lens, Anycast DNS, GRACC/KAPEL, etc
          • Stretched BinderHub platform deployed across UM, MSU, UC, IU, UVic: https://rp1.hl-lhc.io/
            • Want to work with NET2 and SWT2 to meet our milestone of integrating all T2s
        • Bringing Aidan Rosberg, new hire at IU, up to speed
          • Aidan is working on rebuilding several nodes at IU and learning Kubernetes in the process
        • Reana installed on UC AF, working on adding ATLAS IAM auth
        • KAPEL/GRACC integration ongoing
        • Ongoing work to push identity information down into Binder/Jupyter containers - such that we can have POSIX identity + provision storage, etc.
    • 14:25 14:35
      AOB 10m