Conference on Computing in High Energy and Nuclear Physics

Name: Conference on Computing in High Energy and Nuclear Physics
Start: 2024-10-19T08:00:00+02:00
End: 2024-10-25T18:30:00+02:00
Location: No location set

19–25 Oct 2024

Europe/Zurich timezone

Contact Program Chairs

chep2024-pc@cern.ch

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

24 Oct 2024, 14:24

18m

Room 2.A (Seminar Room)

Talk Track 7 - Computing Infrastructure Parallel (Track 7)

Dr David Park (Brookhaven National Laboratory)

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches.

In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes six months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

Yihui Ren (Brookhaven National Laboratory (US)) Dr Ozgur Kilic (Brookhaven National Laboratory) Dr David Park (Brookhaven National Laboratory) Tatiana Korchuganova (University of Pittsburgh (US)) Dr Frederic Suter (Oak Ridge National Laboratory) Joseph Boudreau (University of Pittsburgh (US)) Norbert Podhorszki (Oak Ridge National Laboratory) Paul Nilsson (Brookhaven National Laboratory (US)) Dr Sairam Sri Vatsavai (Brookhaven National Laboratory) Scott Klasky Tasnuva Chowdhury (University of the Witwatersrand (ZA)) Mr Shengyu Feng (Carnegie Mellon University) Verena Ingrid Martinez Outschoorn (University of Massachusetts (US)) Prof. Yiming Yang (Carnegie Mellon University) Tadashi Maeno (Brookhaven National Laboratory (US)) Alexei Klimentov (Brookhaven National Laboratory (US)) Dr Adolfy Hoisie (Brookhaven National Laboratory)

Introspective Dynamic Modelling_CHEP.pdf

Conference on Computing in High Energy and Nuclear Physics

Contact Program Chairs

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

Room 2.A (Seminar Room)

Speaker

Description

Authors

Presentation materials

Choose timezone

Conference on Computing in High Energy and Nuclear Physics

Contact Program Chairs

Speaker

Description

Authors

Presentation materials