31 August 2026 to 4 September 2026
US/Pacific timezone
All in-person registration fee waivers have now been claimed.

Beyond Rounding: Error-Shaping Quantization for Fast Neural Network Inference

2 Sept 2026, 09:30
30m
QI Auditorium

QI Auditorium

Invited Presentation Invited Talks

Speaker

Rayaan Saab (UCSD)

Description

Modern neural networks achieve remarkable performance, but their size and computational cost can make deployment challenging, particularly in settings with tight constraints on memory, latency, or energy. Quantization addresses these challenges by replacing high-precision weights and activations with low-precision representations. The simplest approach is to round each parameter independently, but doing so ignores the considerable redundancy and structure present in modern neural networks.

In this talk, I will describe a line of work that instead views post-training quantization as a sequential error-shaping problem. Rather than discarding the error introduced by each quantization decision, these methods use calibration data to propagate, compensate for, or reshape that error through the remaining degrees of freedom of the network. This perspective leads to efficient algorithms for aggressive low-bit quantization and, importantly, makes aspects of their behavior amenable to rigorous mathematical analysis. I will highlight what this viewpoint reveals about when and why low-precision neural networks can faithfully reproduce their full-precision counterparts.

Presentation materials

There are no materials yet.