Speaker
Description
The CERN Open Data portal provides open access to high-energy physics data for research, education, and outreach. As the volume of hosted data surpasses 5 PB, the need for a sustainable management strategy becomes critical to ensure long-term preservation. Balancing high-performance access for popular datasets with cost-effective storage for rarely accessed data is essential for the continued growth of the repository.
To address these challenges, a cold storage system was moved into production in June 2025. By leveraging tape archives for secondary storage, the portal can preserve massive volumes of data while freeing up primary disk resources. A central feature of this implementation is the self-service staging functionality: unauthenticated users can request the restoration of archived datasets directly through the web interface, with the process handled by an automated queue.
This contribution discusses the integration of cold storage into the Open Data infrastructure and the specific functionalities developed to maintain public access to archived content. Operational insights and usage statistics from the first year of production are also shared, reflecting on how this storage model supports the long-term goals of data preservation in high energy physics.