Speaker
Description
Hosted by the National University of Science and Technology POLITEHNICA Bucharest, the RO-03-UPB site has been an active member of the WLCG computing Grid since 2017 and a member of the ALICE Grid since 2005. Over the course of this collaboration, the site has evolved significantly: originally deployed as a Tier-2 facility, it has grown into a major contributor to the ALICE Grid, currently providing 7,200 CPU cores and 9.6 PB of storage.
To address the ALICE experiment's increased capacity requirements for asynchronous data reconstruction during Long Shutdown 3 and beyond, RO-03-UPB is upgrading to a Tier-1 facility. This contribution outlines the cost-effective strategies employed to achieve this transition.
We discuss the implementation of a disk-based custodial storage solution using EOS in a Redundant Array of Independent Nodes (RAIN) configuration. This approach mitigates the risks associated with disk-based storage while offering distinct advantages over traditional tape-based solutions—specifically, the elimination of complex data staging mechanisms. Consequently, the site can provide a low-latency, high-throughput interface ideal for testing and running asynchronous reconstruction jobs.
Furthermore, to manage the high-volume transfer of raw data between RO-03-UPB and CERN, the site will join the LHCOPN via a 100 Gbps Dense Wavelength Division Multiplexing (DWDM) link. The contribution evaluates this network upgrade alongside alternative solutions. Finally, we address the rigorous availability standards required of a Tier-1 site, detailing the deployment of open-source, industry-standard software for the monitoring, alerting, and backup of Grid services.