Speakers
Description
The Production and Distributed Analysis (PanDA) workload management system was designed with flexibility to adapt to emerging computing technologies in processing, storage, networking, and distributed computing middleware for the global data distribution. PanDA can coordinate processing over heterogeneous computing resources, including dozens of geographically separated high-performance computers. Error occurrence is a common phenomenon in the whole operation that arises through different sources, e.g. human intervention, hardware faults, data transfer issues. Ensuring resilience of the voluminous data is critical during the data life cycle and throughout the execution of scalable workflows to assure that the pathway to viable scientific results is not impeded.
One of the primary goals is to understand, analyze and mitigate error occurrence to ensure the resiliency of the workflow management. We analyzed five months of PanDA metadata related to tasks and jobs. We performed a detailed analysis which gives us a deeper understanding of the error occurrence and patterns and hence developing mitigation strategies. We also categorized all the types of error occurrence. We also studied the impact of such failure on overall system resources, for example. computing hours wastage, memory allocated, etc. We also performed similar analysis at the task metadata as well. Through analysis we developed a time series analysis of the errors and then developed a predictive model for future error occurrence. As a primary challenge, understanding of the field definition is primarily important which is now sparse and not well defined. Even if the definitions are clear, mitigating the interplay of the fields and reducing redundancy is also challenging, especially considering the data volume. Similar challenges exist for the task data as well.