May 27 – 29, 2026
CERN
Europe/Zurich timezone
There is a live webcast for this event.
Ask questions on discord: https://cern.ch/fdf-qa

HGQ: High Granularity Quantization for Real-time Neural Networks and LUT-Based Inference

May 28, 2026, 12:05 PM
30m
500/1-001 - Main Auditorium (CERN)

500/1-001 - Main Auditorium

CERN

400
Show room on map
Talk Algorithm implementation in HDL and HLS AI

Speaker

Chang Sun (California Institute of Technology (US))

Description

Neural networks with sub-microsecond inference latency are required by many critical applications.
Targeting such applications deployed on FPGAs, we present High Granularity Quantization (HGQ), a quantization-aware training framework that optimizes parameter bit-widths through gradient descent.
Unlike conventional methods, HGQ determines the optimal bit-width for each parameter independently, making it suitable for hardware supporting heterogeneous, arbitrary precision arithmetic.
Simultaneously, we introduce HGQ-LUT, a new class of LUT-based layers implemented within HGQ with regular tensor operations during training
, enabling the efficient optimization of LUT-based or hybrid neural networks with more than 2 orders of magnitude faster training speed compared to previous methods.
We show that the HGQ framework achieves superior performance compared to previous arts, achieving significant reduction in resource consumption and latency while maintaining the accuracy.

Talk's Q&A During the talk
Talk duration 20'+10'
Will you be able to present in person? Yes

Author

Chang Sun (California Institute of Technology (US))

Co-authors

Jennifer Ngadiuba (FNAL) Prof. Maria Spiropulu (California Institute of Technology) Qibin Liu (SLAC National Accelerator Laboratory (US)) Thea Aarrestad (ETH Zurich (CH)) Vladimir Loncar (University of Belgrade (RS)) Wayne Luk Zhiqiang Que (Imperial College London)

Presentation materials