Speaker
Description
As the HL-LHC prepares to deliver large volumes of data, the need for an efficient data delivery and transformation service becomes crucial. To address this challenge, a cross-experiment toolset—ServiceX—was developed to link the centrally produced datasets to flexible, user-level analysis workflows. Modern analysis tools such as Coffea benefit from ServiceX as the first step in event selection, efficiently reducing file sizes through a Kubernetes infrastructure. ServiceX allows query-based data transformers with different backends that provide remote access to tuple ROOT and parquet files and experiment-specific EDM files (e.g. ATLAS xAOD). By facilitating the transformation of heterogeneous data formats into columnar representations, ServiceX can reduce data loads and accelerate analyses in a user-friendly way.
This talk will discuss the ServiceX infrastructure, which enables remote querying and data skimming in distributed systems. Additionally, we will discuss the range of analysis workflows in which ServiceX can be integrated and the recent developments that aim to extend it, notably with its integration in standard ATLAS analysis frameworks. Such developments also introduce new use-case features that increase the available ServiceX-based toolset to analyzers, ensuring the integration of ServiceX in different steps of an analyser workflow.
Significance
ServiceX's deployment is broadening with the inclusion of standard analysis frameworks in the backend. Specific functionalities for introspecting datasets are released, and other additions are planned, giving analysts more useful tools.
References
https://indico.cern.ch/event/1330797/contributions/5796587/
| Experiment context, if any | ATLAS/CMS physics data analysis |
|---|