8โ€“12 Sept 2025
Hamburg, Germany
Europe/Berlin timezone

Error Analysis of PanDA metadata

10 Sept 2025, 11:00
30m
ESA W 'West Wing'

ESA W 'West Wing'

Poster Track 1: Computing Technology for Physics Research Poster session with coffee break

Speakers

Raees Ahmad Khan (University of Pittsburgh (US))Ms Tania Korchuganova

Description

The Production and Distributed Analysis (PanDA) workload management system was designed with flexibility to adapt to emerging computing technologies in processing, storage, networking, and distributed computing middleware for the global data distribution. PanDA can coordinate processing over heterogeneous computing resources, including dozens of geographically separated high-performance computers. Error occurrence is a common phenomenon in the whole operation that arises through different sources, e.g. human intervention, hardware faults, data transfer issues. Ensuring resilience of the voluminous data is critical during the data life cycle and throughout the execution of scalable workflows to assure that the pathway to viable scientific results is not impeded.
One of the primary goals is to understand, analyze and mitigate error occurrence to ensure the resiliency of the workflow management. We analyzed five months of PanDA metadata related to tasks and jobs. We performed a detailed analysis which gives us a deeper understanding of the error occurrence and patterns and hence developing mitigation strategies. We also categorized all the types of error occurrence. We also studied the impact of such failure on overall system resources, for example. computing hours wastage, memory allocated, etc. We also performed similar analysis at the task metadata as well. Through analysis we developed a time series analysis of the errors and then developed a predictive model for future error occurrence. As a primary challenge, understanding of the field definition is primarily important which is now sparse and not well defined. Even if the definitions are clear, mitigating the interplay of the fields and reducing redundancy is also challenging, especially considering the data volume. Similar challenges exist for the task data as well.

Authors

Paul Nilsson (Brookhaven National Laboratory (US)) Sankha Baran Dutta (Brookhaven National Laboratory (US)) Ms Tania Korchuganova

Co-authors

Adolfy Hoisie (Brookhaven National Laboratory (US)) Alexei Klimentov (Brookhaven National Laboratory (US)) David Park (Brookhaven National Laboratory) Fatih Furkan Akman (University of Massachusetts (US)) Frederic Suter John Rembrandt Steele (University of Massachusetts (US)) Joseph Boudreau (University of Pittsburgh) Kuan-Chieh Hsu (Brookhaven National Laboratory (US)) Norbert Podhorszki (Oak Ridge National Laboratory) Ozgur Ozan Kilic (Brookhaven National Laboratory) Raees Ahmad Khan (University of Pittsburgh (US)) Ray Ren (Brookhaven National Laboratory (US)) Sairam Sri Vatsavai (Brookhaven National Laboratory (US)) Scott Klasky Mr Shengyu Feng Shinjae Yoo Tadashi Maeno (Brookhaven National Laboratory (US)) Dr Tasnuva Chowdhury (Brookhaven National Laboratory (US)) Verena Ingrid Martinez Outschoorn (Fermi National Accelerator Lab. (US)) Wei Yang (SLAC National Accelerator Laboratory (US)) Prof. Yiming Yang

Presentation materials