Speaker
Description
Author: Andreea Prigoreanu (University Politechnica Bucharest)
on behalf of the ALICE collaboration
The processing of ALICE experiment data relies on high-quality and reliable storage. The central file catalogue serves as the database that tracks over 2.6 billion files and their locations across more than 50 storage elements on the ALICE Grid. It is essential that the physical storage contents remain consistent with the catalogue to prevent errors when users or jobs access data and to avoid data loss. One of the solutions ALICE currently uses to monitor the consistency is a distributed file crawler that periodically evaluates samples of files from each storage element to gather statistics about the number of corrupted or inaccessible files. While effective for detecting inconsistencies, the method is constrained by its sampling-based nature, and addressing the issues requires manual investigation and intervention on the storage itself.
EOS is the most widely deployed storage system across the ALICE disk-based storage sites, managing approximately 275 PB of the total 360 PB of ALICE disk-based data. By leveraging the powerful internal checking tools provided by EOS, the ALICE-wide EOS Integrity and Recovery System is a new solution to address data integrity for files stored on EOS instances throughout the ALICE Grid. It complements the existing strategy by collecting and analyzing error information from EOS FSCK reports, accessed through the HTTP interface available in recent EOS versions. This approach enhances data monitoring across ALICE storage systems, providing a comprehensive view of storage health for EOS deployments. In addition to identifying issues and generating file consistency reports, the Integrity and Recovery System automates the recovery process by invoking the experiment's recovery procedures to reconcile the contents of the storages with the central file catalogue.
I will present the system's key aspects: storage configuration requirements, automated FSCK report retrieval and analysis, the decision algorithm for flagging files requiring recovery, and the integration with the recovery procedures.