15–19 Sept 2025
CERN
Europe/Zurich timezone

GPU (de)compression

16 Sept 2025, 15:45
5m
40/S2-A01 - Salle Anderson (CERN)

40/S2-A01 - Salle Anderson

CERN

95
Show room on map
2. Optimal AI deployment for Online Data Processing Optimal AI deployment for Online Data Processing

Speaker

Jolly Chen (CERN & University of Twente (NL))

Description

In the HL-LHC era, we expect a significant increase in the amount of data and compression can serve as an effective tool to reduce storage requirements. As more and more computing facilities are equipped with GPUs, enabling lossless (de)compression directly on GPUs can reduce costly memory transfers to/from the GPU, remove reliance on the CPU for compression tasks, and increase the effective storage capacity of GPU memory. These advantages could greatly improve the performance in applications where the GPU data transfers or memory capacity become the bottleneck. However, current GPU compressors are either too slow or do not compress well enough for our purposes.

A compression algorithm essentially consists of a pipeline of encoders. We propose to train an ML model that can select the most optimal combination of GPU encoders, which builds a “new” compression algorithm that optimises the tradeoff between throughput and compression ratio for HEP data and HEP workflows. With this project, we aim to determine how much we benefit from GPU (de)compression in practice or whether existing encoding solutions are unsuitable for our needs, and we need to develop an in-house encoding solution tailored to our data

Further information can be found here: https://docs.google.com/document/d/1hU86jlNSCHE-pV8m-CNQ3dj_fo2RfpYOzJf6lsWxrA8/edit?tab=t.0

CERN group/ Experiment

NGT 1.7 and EP-SFT

Working area Area 2: Optimal AI deployment for Online Data Processing
Project goals Problem: inefficient GPU workflows due to memory transfer bottlenecks and/or insufficient GPU memory Intermediate goal: create a ML model that can select the optimal combination of encoders (i.e., GPU compression algorithm) for a given workflow and dataset, optimizing the tradeoff between compression throughput and compression ratio. Final Goal: determine how much performance can be gained from GPU (de)compression through smaller data transfers and/or larger effective GPU memory storage by storing compressed data on the GPU and operating on GPU decompressed chunks.
Timeline Year 1: Foundations & First Prototype. - Familiarize with LC framework, GPU compression libraries, and HEP workflows. - Gather representative HEP datasets and establish a benchmark suite. - Define performance metrics (compression ratio, throughput, transfer latency, CPU/GPU utilization). - Survey encoder candidates (existing GPU compressors + CPU baselines). - Deliverable: baseline performance study + first ML-based prototype (approach 1). Year 2: Expansion & Validation. - Second prototype (approach 2). - Run systematic benchmarking across multiple datasets and workflows. - Compare two approaches, highlight trade-offs. - Deliverable: comparative study + recommendation of promising direction. Year 3: Optimization & Integration - Optimize the most promising prototype (scalability, integration into HEP frameworks, GPU memory handling). - Large-scale validation on production-like HEP workloads. - Explore fallback path: feasibility study for in-house GPU encoder if existing components are insufficient. - Deliverable: model/software release, integration guidelines, comprehensive report.
Available person power 0
Additional person power request 1 GRAD/DOCT + 0.2 STAFF (supervision)
Is this an already ongoing activity? No
Indicative hardware resources needs server-grade GPU access

Authors

Jakob Blomer (CERN) Jolly Chen (CERN & University of Twente (NL)) Sanjiban Sengupta (CERN, The University of Manchester)

Presentation materials

There are no materials yet.