Speaker
Description
The ALICE Event Processing Nodes (EPN) farm is a high-density GPU HPC system designed primarily for real-time reconstruction of 50 kHz Pb-Pb collisions during LHC Run 3. It is the largest computer farm at CERN in terms of compute capacity. Comprising 350 nodes and 2800 GPUs, with a peak performance of ~42 PFLOP/s single precision, the HPC infrastructure has been operated throughout Run 3 by a dedicated team of 2 to 3 individuals at a time. This contribution presents the organisational, technical, and architectural choices that enabled this 24/7-supported, high-reliability, low-maintenance operational model.
The team operates the full stack of the HPC environment: electrical and cooling systems, networking, servers and GPUs, firmware management and orchestration layers. Automation of provisioning, configuration management, monitoring, and recovery procedures ensures reliable operations with minimal manual intervention. The software stack: OS and driver management, InfiniBand-based high-throughput data distribution,orchestration tools, and integration with experiment and detector control systems; all provide a base for reliable operations and a separation between online and asynchronous processing modes, and maintenance and testing environments. These design principles enable sustainable operation by reducing manpower requirements and optimizing hardware utilization.
The talk summarizes lessons learned from several years of continuous operation of a physics-critical, high-throughput, GPU-accelerated HPC facility and highlights principles applicable to other large-scale scientific computing facilities aiming for sustainability and low operational overhead.