Speaker
Description
Transformers are promising for real-time intelligent systems, but their arithmetic complexity makes microsecond-scale FPGA inference challenging. This paper presents a LUT-based Transformer framework that combines heterogeneous quantization and LUT-aware training to map compact Transformer models into hardware-efficient lookup-table structures. The proposed method jointly optimizes precision, arithmetic structure, and FPGA resource usage, enabling fine-grained mixed-precision inference with reduced logic cost and latency. Compared with Linformer and compact Transformer baselines, our LUT-based Transformer achieves better accuracy–latency–resource trade-offs and delivers microsecond-scale FPGA inference. These results show that LUT-aware mixed-precision training is an effective approach for deploying Transformers in latency-critical applications.