25–29 May 2026
Chulalongkorn University
Asia/Bangkok timezone

Evolving PanDA: Toward Sustainable, Intelligent, and Heterogeneous Workload Management in ATLAS and beyond

26 May 2026, 14:03
18m
MHMK 202

MHMK 202

Oral Presentation Track 4 - Distributed computing Track 4 - Distributed computing

Speaker

Xin Zhao (Brookhaven National Laboratory (US))

Description

The ATLAS experiment at the CERN Large Hadron Collider relies on a worldwide distributed computing infrastructure to process millions of production and analysis jobs daily across grid, cloud, and HPC resources. The ATLAS Distributed Computing (ADC) system integrates workload, data, and resource management services to ensure efficient use of heterogeneous environments. Within ADC, the PanDA Workload Management System (WMS) provides large-scale job brokerage, pilot submission, and monitoring. Recent development extends PanDA to address sustainability, hardware diversity, and intelligent automation. A new module estimates per-job CO₂-equivalent emissions, combining runtime metadata with regional carbon-intensity data. The resulting gCO₂ values are stored in the PanDA database and visualized through monitoring dashboards to raise awareness of computing-related emissions. The PanDA brokerage has been extended to support GPU-based scheduling, using a redesigned JSON resource description that encodes CUDA version, GPU model, memory, and benchmark data to enable precise resource matching. The Worker Node Map collects detailed CPU and GPU specifications reported by pilots and correlates them with HEPiX benchmark results to improve CPU-time normalization and provide operational insight into site homogeneity and hardware aging. Ask PanDA introduces an AI-driven assistant that orchestrates multiple specialized clients in a coordinated workflow and employs retrieval-augmented generation to provide contextual answers on PanDA operations. These developments represent a significant step toward a more transparent, efficient, and sustainable computing ecosystem for ATLAS and future large-scale scientific workflows.

Authors

Aleksandr Alekseev (The University of Texas at Arlington (UTA)) Ben Bruers (Deutsches Elektronen-Synchrotron (DE)) Dr Edward Karavakis (Brookhaven National Laboratory (US)) Fa-Hui Lin (University of Texas at Arlington (US)) Fernando Harald Barreiro Megino (University of Texas at Arlington) Kaushik De (University of Texas at Arlington (US)) Mikhail Borodin (CERN) Misha Borodin (University of Texas at Arlington (US)) Paul Nilsson (Brookhaven National Laboratory (US)) Rodney Walker (Ludwig Maximilians Universitat (DE)) Tadashi Maeno (Brookhaven National Laboratory (US)) Tatiana Korchuganova (University of Pittsburgh (US)) Wen Guan (Brookhaven National Laboratory (US)) Xin Zhao (Brookhaven National Laboratory (US))

Presentation materials