Speaker
Description
Machine learning models used in real-time and resource-constrained environments, such as hardware triggers, online reconstruction pipelines, and FPGA/GPU inference systems, must satisfy strict latency, memory, and numerical precision requirements. Achieving these targets typically requires extensive tuning of training schedules, quantization settings, sparsity levels, and architectural parameters. In current workflows, this optimization process is often manual and difficult to reproduce, especially when multiple objectives (e.g., accuracy, latency, and on-chip footprint) must be balanced simultaneously.
To address this challenge, we introduce a new hyperparameter optimization (HPO) platform within the PQuantML library, developed as part of the Next-Generation Trigger (NGT) project. The platform provides an integrated framework for automated exploration of compression parameters and fine-tuning strategies, built on Optuna for adaptive sampling and MLflow for experiment tracking. Users define search spaces and evaluation metrics through configuration files, enabling large-scale optimization experiments without modifying model code. The system supports Bayesian/TPE sampling, early-pruning strategies, and multi-objective optimization, allowing the search to target both physics performance metrics and hardware-level constraints.
The module is designed for distributed execution on the NGT cluster, enabling hundreds of parallel trials to evaluate trade-offs between accuracy, sparsity, bit precision, and latency. We demonstrate the framework on representative convolutional and classifier models used in real-time ML studies, showing how automated optimization systematically identifies the best configurations that meet HL-LHC latency and resource budgets while maintaining high physics performance. This integrated HPO capability strengthens PQuantML as a toolchain for preparing deployable ML models and provides a reproducible workflow for tuning models destined for the NGT and online computing systems.