Speaker
Description
Neural networks with sub-microsecond inference latency are required by many critical applications.
Targeting such applications deployed on FPGAs, we present High Granularity Quantization (HGQ), a quantization-aware training framework that optimizes parameter bit-widths through gradient descent.
Unlike conventional methods, HGQ determines the optimal bit-width for each parameter independently, making it suitable for hardware supporting heterogeneous, arbitrary precision arithmetic.
Simultaneously, we introduce HGQ-LUT, a new class of LUT-based layers implemented within HGQ with regular tensor operations during training
, enabling the efficient optimization of LUT-based or hybrid neural networks with more than 2 orders of magnitude faster training speed compared to previous methods.
We show that the HGQ framework achieves superior performance compared to previous arts, achieving significant reduction in resource consumption and latency while maintaining the accuracy.
| Talk's Q&A | During the talk |
|---|---|
| Talk duration | 20'+10' |
| Will you be able to present in person? | Yes |