31 August 2026 to 4 September 2026
US/Pacific timezone
All in-person registration fee waivers have now been claimed.

XFMT: LUT-based Transformer for microsecond-scale inference

2 Sept 2026, 11:24
12m
QI Auditorium

QI Auditorium

Presentation Contributed Talks

Speaker

Arianna Cox (Imperial College (GB))

Description

Transformers are promising for real-time intelligent systems, but their arithmetic complexity makes microsecond-scale FPGA inference challenging. This paper presents a LUT-based Transformer framework that combines heterogeneous quantization and LUT-aware training to map compact Transformer models into hardware-efficient lookup-table structures. The proposed method jointly optimizes precision, arithmetic structure, and FPGA resource usage, enabling fine-grained mixed-precision inference with reduced logic cost and latency. Compared with Linformer and compact Transformer baselines, our LUT-based Transformer achieves better accuracy–latency–resource trade-offs and delivers microsecond-scale FPGA inference. These results show that LUT-aware mixed-precision training is an effective approach for deploying Transformers in latency-critical applications.

Authors

Chang Sun (California Institute of Technology (US)) Arianna Cox (Imperial College (GB)) Lauri Antti Olavi Laatu (Imperial College (GB)) Benedikt Maier (Imperial College (GB)) Aadeesh Sharma Leo Rozanov Wayne Luk Alex Tapper (Imperial College London) Prof. Maria Spiropulu (California Institute of Technology) Zhiqiang Que (University of Bristol)

Presentation materials

There are no materials yet.