US ATLAS Computing Facility (Possible Topical)

US/Eastern
Description

Facilities Team Google Drive Folder

Zoom information

Meeting ID:  993 2967 7148

Meeting password: 452400

Invite link:  https://umich.zoom.us/j/99329677148

 

 

    • 13:00 13:05
      WBS 2.3 Facility Management News 5m
      Speakers: Alexei Klimentov (Brookhaven National Laboratory (US)), Dr Shawn Mc Kee (University of Michigan (US))

      Thanks for getting the quarterly reporting done.

      We are awaiting a summary of the scrubbing

        - Once we have that, we will need to review and finalize new milestones proposed at scrubbing

        - We will add the approved milestones to the new Jul-Sep milestone spreadsheet 

      Lots of items pending beyond scrubbing outcomes:

        - Awaiting end-of-CA funds and the resulting procurement processes

        - Preparing for DC27: we need to discuss USATLAS plans for mini-capacity and mini-capability challenges

        - Fall events:  IRIS-HEP retreat, LHCOPN/LHCONE co-located with 5th Global Research Platform meeting, HEPiX in Lincoln, NE, HSF/WLCG in early November, SC26 in Chicago

      Other items?

       

       

      Quick recap

      This meeting was a facility coordination meeting where team members provided updates on various operational aspects of their systems. Shawn opened the meeting by thanking everyone for completing quarterly reporting and noting they were waiting for results from the scrubbing process, which would determine new milestone reviews and approvals. Brian reported that Crick and Topology contact integration had been completed with API key support, and XRootD 6.0/6.1 was in testing, while also mentioning recent HTCondor security releases that affected FS authentication in certain configurations. Thomas provided updates on tier one operations, noting smooth running conditions and work on Condor configuration rewrites, while Carlos reported that DCACHE and HPSS operations remained stable with some recent interventions for storage pool rack reorganization and certificate updates. Frederick mentioned that tier two sites were running reasonably well with some minor glitches, and expressed cautious optimism about end-of-CA funding despite time constraints. Rui updated on HPCs, noting that Eris was online and SQL server migration was proceeding, while Fengping discussed DNS query optimization issues and GPU scheduling challenges at Chicago. Qiulan reported on Jupyter environment support for BLR championship and user data storage analysis showing most users accessing local group and scratch group data rather than their personal areas. Ivan concluded with updates on CRIC data synchronization, tape storage cleanup, and preparation for long periods without data taking, also mentioning out-of-memory issues at Southwest Tier 2 and problems with improperly configured task requests.

      Next steps

      Brian

      • Coordinate with Ilya to test XRootD 6.0 and 6.1 in the Pelican configuration.
      • Tighten up testing procedures for HTCondor security fix, especially for sites using FS authentication with Natted or multi-homed hosts.

      Fengping

      • Investigate and optimize scheduling for GPU nodes to prevent CPU exhaustion blocking GPU scheduling.

      Gamboa

      • Deploy DCache 11.2.05/06 and enable CyTACs, targeting around September.
      • Continue action plan for dual home pool nodes, using hardware release cycles or identifying budget for full conversion.

      Ivan

      • Coordinate with CRIC developers to propagate API changes to ATLAS sites.
      • Investigate out-of-memory issue at Southwest Tier 2 with XRootD developers, focusing on HERIC certificate authority revocation list files.
      • Propose to Panda to limit scouting to one task per request initially to prevent mass job failures.

      Qiulan

      • Continue data access analysis for user data storage and further check user usage patterns.

      Rui

      • Follow up with supporting team on new review for switching to one niche instead of SQL server for AICF site.

      Shawn

      • Review and finalize new milestones based on the scrubbing outcome summary once received from Paolo and Verena.

      Suzanne

      • Produce the new July through September milestone spreadsheet once the scrubbing outcome summary is available.

      Thomas

      • Continue Condor config rewrites for all pools and plan for Alma 9.7 to 9.8 update for the shared pool.

      Summary

      Aging Calculation Project Discussion

      Frederick offered to help Shigeki with aging calculation, mentioning his expertise in this area. Alexei discussed the need for uniform aging tables and mentioned sending Shigeki a spreadsheet as an example. Frederick noted that the project had slowed down after Chris Halliwell's departure and suggested they could now recover and continue the work. Shawn briefly mentioned that they were up to five items on the agenda, hoping to complete the meeting more quickly.

      Project Scrubbing and Milestone Review

      Shawn announced that he needs to discuss scrubbing outcomes with Paolo and Verena after the meeting, as they are working on providing a summary of the results. Once the scrubbing information is available, the team will review and finalize new milestones that may need modification based on potential changes from the scrubbing process. The team also discussed upcoming tasks including preparing for DC27, which is scheduled for February-March 2027, and various fall events including Iris Hip retreat, LHCP meetings, HEPICS, and Supercomputing 2026 in Chicago.

      Digital Discussion Summary

      The transcript appears to be a series of fragmented questions or statements in Russian, likely related to a discussion about digital topics, companies, and percentages. Due to the fragmented nature of the transcript and lack of clear context or decisions made, it is not possible to provide a meaningful summary that captures decisions, alignments, or next steps.

      DOE Project Funding Discussion

      Shawn confirmed that DOE is willing to back a project, though there were concerns raised about resource allocation and expectations for volunteer work without proper compensation. The discussion included questions about accessing American Science Cloud and obtaining tokens, though these details remained unclear. The conversation ended with no additional major topics raised by participants.

      OSG LHC Updates and Releases

      Brian provided updates on OSG LHC, including the completion of Crick and Topology contact integration with API key support and the availability of XRootD 6.0 and 6.1 for testing. He highlighted recent HTCondor security releases affecting FS authentication, particularly for Natted or multi-homed hosts, and requested feedback from the team on any issues. Brian also announced plans for the OSG and HTCondor Pelican 26 release in September, which will mark the end of life for OSG and HTCondor 24.

      Condor Version Update Progress Report

      Thomas reported that tier one operations have been smooth for the past couple of weeks with no issues or problems. He has been working on updating the Condor version on tier three submit hosts and translating job transforms when moving from earlier LTS version 25 to 25.11.1. Thomas mentioned that future work will include rebuilding the pool for an Alma 9.7 to 9.8 update, though the date for this has not been determined.

      Storage Operations Status Update

      Gamboa reported that DCACHE and HPSS operations remained stable, with successful completion of storage pool rack reorganization in July that required shutting down five nodes for cable reorganization. The team completed updates to puppet classes to support new CA certificates, with all ATLAS storage elements now consuming certificates signed by the new CA. Gamboa also announced the start of a controlled deletion campaign for tape files, targeting 5 million files with 14% complete, expecting to finish by the end of the week.

      CyTags DCache Implementation Timeline Updates

      Gamboa provided updates on the timeline for BNL to implement CyTags-enabled DCache. The new release 11.2.05 is expected to be deployed around September, with a potential new release 11.2.06 coming in two weeks. Regarding pool nodes in the DMZ, Gamboa explained that while some pools with dual home capability are already part of the production system, full implementation requires budget allocation for hardware release cycles, which is currently in discussion.

      System Status and Implementation Updates

      Gamboa discussed the utility of a current system beyond BNL and inquired about its status at other sites, with Shawn indicating limited knowledge of other sites requiring similar functionality. The discussion highlighted challenges with proxy mode implementation, particularly for third-party copy and DCASH architecture. Frederick provided an update on tier two sites, noting reasonable running conditions over the past two weeks despite some glitches, and expressed cautious optimism about end-of-funding, pending scrubbing results.

      HPC System Updates and Issues

      Rui reported that HPC permits remain under maintenance with an extended timeline beyond next Monday, while Eris is online but no updates on top up have been received. The SQL server passed network and security review and will proceed to step three, with additional testing planned for one niche instead of squared AICF site. Fengping identified two issues: excessive DNS queries due to CoreDNS configuration and CPU allocation challenges on GPU nodes affecting scheduling, particularly from charging servers and ServiceX transformers. Shawn confirmed the CPU and GPU co-scheduling requirement and acknowledged the need for further investigation into optimization possibilities.

      BLR Analysis and Operations Updates

      Qiulan reported on BLR analysis for safety, including updates on the Jupyter environment and user data storage analysis showing most users access local groups rather than their personal data areas. Ivan discussed continuous integration and operations updates, including data synchronization issues, tape storage recovery, and preparations for the long period without data taking, along with addressing out-of-memory problems at Southwest Tier 2 sites and proposing to limit Panda requests to one task per request. The conversation ended with announcements about upcoming Facility Coordination and regular meetings.



    • 13:05 13:10
      OSG-LHC 5m
      Speakers: Brian Hua Lin (University of Wisconsin), Matyas Selmeci
      • CRIC + Topology contact integration complete
      • XRootD 6.0 and 6.1 available in upcoming-testing
      • Recent HTCondor secruity releases affecting FS authentication: known issue with NAT'ed / multi-homed hosts
      • Starting to plan OSG / HTCondor / Pelican 26
        • Aiming for a September release
        • OSG / HTCondor 24 will be EOL'ed upon release of 26
    • 13:10 13:30
      WBS 2.3.1: Tier1 Center
      Convener: Alexei Klimentov (Brookhaven National Laboratory (US))
      • 13:10
        Tier-1 Infrastructure 5m
        Speaker: Jason Smith
      • 13:15
        Compute Farm 5m
        Speaker: Thomas Eric Smith (Brookhaven National Laboratory (US))

        Tier 1 is operating smoothly, no issues or changes in last week

        • Tier 1 pool upgrade from Almalinux 9.7 to 9.8 date TBD
        • Tier 1 pool already has condor config refactor, so OS upgrade should be simple

         

        Tier 3 / AF (members of the shared pool)

        • Work being finished on the condor config rewrites, testing is underway. This pool is a bit more complex than the T-1
        • attsubXX hosts (atlas T3 /AF interactive submit hosts) upgraded to condor v25.11.1 for security updates and features
          • Old job transform syntax went EOL during the 25.X feature series, needed to update legacy transforms (this is very easy to do)
          • Sites should make sure they have updated all their transforms to the new syntax before moving to future condor 26.0 LTS or 25.X feature releases
      • 13:20
        Storage 5m
        Speakers: Carlos Fernando Gamboa (Brookhaven National Laboratory (US)), Carlos Fernando Gamboa (Department of Physics-Brookhaven National Laboratory (BNL)-Unkno)

         

        • Stable and reliable dCache and HPSS operations throughout the reporting period.
        • Completed storage pool infrastructure maintenance (rack cable reorganization) on dc252–dc256 on July 21, 2026, from 9:00 a.m. to 12:00 p.m.
        • Successfully deployed and validated updated certificates across all ATLAS dCache components to support InCommon Generation 4, included in IGTF distribution v1.143 (released June 22, 2026).
        • Continued the internal HPSS tape deletion campaign, which is now approximately 14% (700k files) to complete.
      • 13:25
        Tier1 Operations and Monitoring 5m
        Speaker: Ofer Rind (Brookhaven National Laboratory)
    • 13:30 13:40
      WBS 2.3.2 Tier2 Centers

      Updates on US Tier-2 centers

      Conveners: Fred Luehring (Indiana University (US)), Rafael Coelho Lopes De Sa (University of Massachusetts (US))
      • Reasonable running over the last two weeks.
        • Various minor reductions in production:
          • Security updates still causing trouble. 
          • AGLT2 had configuration error causing issues. Reverted.
          • MWT2 Illinois preventive maintenance and IU power work.
          • NET2 back off error causing trouble.
          • OU Had disk fill because cleaning restart script did not run.
          • CPB had trouble yesterday with a bad job set.
      • Quarterly reporting in.
      • I am cautiously optimistic on the end of CA funding.
        • Need  to see the scrubbing results before proceeding
        • The amounts will probably differ on some level from what we have discussed.
        • I hope that I have more solid news soon.
    • 13:40 13:50
      WBS 2.3.3 Heterogenous Integration and Operations

      HIOPS

      Convener: Rui Wang (Argonne National Laboratory (US))
      • 13:40
        HPC Operations 5m
        Speaker: Rui Wang (Argonne National Laboratory (US))

        Perlmutter: under maintenance (extended beyond Aug 3)

        • IRIS is online, waiting for top-up

        TACC:

        • Squid passed network and security review. They are happy to migrate and maintain it for Stampede3
        • Start review process for Varnish

        ALCF:

        • Helped on Wen Guan's account setup at ALCF and have him added to our DD allocation to test IRI harvest on Crux
        • Local FullSim job succeeded with CVMFSexec and local Squid 
      • 13:45
        Integration of Complex Workflows on Heterogeneous Resources 5m
        Speaker: Doug Benjamin (Brookhaven National Laboratory (US))
    • 13:50 14:10
      WBS 2.3.4 Analysis Facilities
      Convener: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 13:50
        Analysis Facilities - BNL 5m
        Speaker: Qiulan Huang (Brookhaven National Laboratory (US))
      • 13:55
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:00
        Analysis Facilities - Chicago 5m
        Speaker: Fengping Hu (University of Chicago (US))

        K8s optimization

        • Excessive wasted external DNS queries traced to the default DNS configuration (ndots:5). Mitigations under consideration:
          • Lower ndots via admission policies
          • Add an explicit dnsConfig to pods
          • Use trailing dots in applications to bypass the search path entirely

        Scheduling challenge observed

        • A GPU node exhausted its CPU allocation, blocking GPU jobs from being scheduled.
        • Contributing factors: a Triton inference server with a large CPU request, combined with bursts from ServiceX transformers.
        • Priority and preemption policies may be worth investigating as a longer-term fix.
    • 14:10 14:30
      WBS 2.3.5 Continuous Operations
      Conveners: Ivan Glushkov (Brookhaven National Laboratory (US)), Ofer Rind (Brookhaven National Laboratory)
      • Security: The OSG admin and security contacts are being propagated properly from OSG topology to ATLAS CRIC RC sites correctly now (CRIC-421)
      • ADC has discovered a historical replication rule for data on Tier1 tape and once ADC has remove this they have cleaned up the data which was wrongly replicated on tape. This results in 10 PB drop in tape utilization (ATLDDMOPS-5849)
      • ADC is preparing for long period without data takeing by reallocating T0 resources (disk and compute)
      • 14:10
        ADC Operations, US Cloud Operations: Site Issues, Tickets & ADC Ops News 5m
        Speaker: Kaushik De (University of Texas at Arlington (US))
        • XRootD OOM issue that SWT2 has been struggling with since the 20th - it looks like other sites (AU, UK etc) are seeing similar things. We are talking to them and Andy/Wei. The problem is quite severe for SWT2.
        • Shifters and SWT2 site people found some badly defined tasks last week (max allowed memory errors). DPA was notified.
      • 14:15
        Services DevOps 5m
        Speaker: Ilija Vukotic (University of Chicago (US))
      • 14:20
        Facility R&D 5m
        Speaker: Robert William Gardner Jr (University of Chicago (US))
      • 14:25
        Cybersecurity plan(s) 5m
        Speakers: Robert William Gardner Jr (University of Chicago (US)), Shigeki Misawa (Brookhaven National Laboratory (US))
    • 14:30 14:40
      AOB 10m