Speaker
Description
The rising computational demands of growing data rates and complex machine learning (ML) algorithms in large-scale scientific experiments have driven the adoption of the Services for Optimized Network Inference on Coprocessors (SONIC) framework. SONIC accelerates ML inference by offloading tasks to local or remote coprocessors, optimizing resource utilization. Its portability across diverse coprocessors enhances data processing and model deployment efficiency for advanced research in high-energy physics (HEP) and multi-messenger astrophysics (MMA). We developed SuperSONIC, a scalable server infrastructure for SONIC, enabling the deployment of computationally intensive tasks on Kubernetes clusters equipped with graphics processing units (GPUs). Leveraging NVIDIA’s Triton Inference Server, SuperSONIC decouples client workflows from server infrastructure, standardizing communication, improving throughput, and enabling robust load balancing and monitoring. Successfully deployed for the CMS and ATLAS experiments at CERN’s Large Hadron Collider, the IceCube Neutrino Observatory, and the LIGO gravitational-wave observatory, SuperSONIC has been tested on Kubernetes clusters at Purdue University, the National Research Platform, and the University of Chicago. SuperSONIC provides a reusable, configurable framework that addresses Cloud-native challenges, enhancing accelerator-based inference efficiency across diverse scientific and industrial applications.