May 25 – 29, 2026
Chulalongkorn University
Asia/Bangkok timezone

Design and Implementation of an AI-Driven Disk Failure Prediction System for Large-Scale HEP Storage Clusters

May 27, 2026, 2:57 PM
18m
MHMK M01

MHMK M01

Oral Presentation Track 1 - Data and metadata organization, management and access Track 1 - Data and metadata organization, management and access

Speaker

LI Haibo lihaibo

Description

In high energy physics (HEP) experiments, large-scale storage clusters typically comprise tens of thousands of disks, and their reliability is essential for continuous data acquisition, processing, and long-term preservation. Traditional rule-based disk failure detection approaches are increasingly insufficient for such environments due to heterogeneous device types, complex workload patterns, and dynamically changing operational behaviors. To address these challenges, this paper presents the design and implementation of an AI-driven disk failure prediction system tailored for large-scale HEP storage clusters.

The system leverages SMART low-level telemetry and introduces a quasi-online feature selection mechanism capable of automatically identifying key indicators strongly correlated with disk failures, thereby enabling lightweight and scalable online feature updates. By combining historical failure statistics with real-time operational metrics through a multidimensional threshold fusion strategy, we develop an intelligent prediction model capable of forecasting disk failures on a daily basis. Deployment in a production environment with over ten thousand disks demonstrates that the system significantly improves early failure detection and reduces unplanned storage downtime.

The system adopts a microservice architecture with standardized RESTful APIs, supports elastic deployment via Docker and Kubernetes, and integrates seamlessly with Prometheus-based monitoring stacks. These capabilities enable automated inference, anomaly alerting, and system state visualization. Experimental results show that the proposed AI-based disk failure prediction system achieves strong prediction accuracy, real-time responsiveness, and scalability, providing a practical and effective solution for enhancing reliability in future exabyte-scale HEP storage infrastructure.

Authors

LI Haibo lihaibo Siqi Hou Dr Yaodong CHENG (Institute of High Energy Physics, Chinese Academy of Sciences) Dr Yujiang BI (Institute of High Energy Physics, Chinese Academy of Sciences) Mr jiangang zhang

Presentation materials