31 August 2026 to 4 September 2026
US/Pacific timezone

Benchmarking BitNet-Style Low-Precision Networks for Fast Jet Classification Toward FPGA Deployment

31 Aug 2026, 17:30
1h 30m
QI Courtyard

QI Courtyard

Poster Extended Abstract Reception & Poster Session

Speaker

Dorian Sloot (Austrian Academy of Sciences (AT))

Description

Low-latency machine learning is essential for real-time decision systems in high-energy physics, where strict resource and latency constraints limit deployable model complexity. Aggressive quantization can potentially reduce memory and arithmetic requirements, but its effect on predictive performance must be understood before deployment. We present a controlled comparison of low-precision neural networks for five-class jet classification, focusing on BitNet-style binary- and ternary-weight multilayer perceptrons using the BitLinear implementations from the BitHEP codebase. We compare these networks with established quantization-aware approaches and prepare conversion artifacts for subsequent FPGA evaluation.

We use the public OpenML hls4ml_lhc_jets_hlf dataset, containing 830,000 jets represented by 16 high-level jet-substructure observables and balanced across gluon, light-quark, W, Z, and top classes. Although this is not a trigger-native dataset, its compact tabular representation provides a useful test case for constrained real-time inference. All models use fixed stratified training, validation, and test partitions, identical preprocessing, and a common evaluation procedure. We compare floating-point multilayer perceptrons, QKeras quantization-aware networks, heterogeneous granularity quantization, and BitNet-style binary-weight and 1.58-bit ternary-weight networks, retaining method-appropriate training configurations where required.

Across three training seeds, the floating-point baseline achieved 76.64% accuracy and a macro AUC of 0.9421. Topology-matched 7–12-bit QKeras models closely reproduced this performance. The BitNet-1.58 model achieved 74.21% accuracy and consistently outperformed the tested binary-weight BitNet variants, which achieved 71.98–72.28%. At a nominal 100 kHz q/g-background proxy operating point, the floating-point and 7–12-bit QKeras models retained approximately 45% combined W, Z, and top efficiency, compared with 35.9% for BitNet-1.58 and 29.7–33.3% for the binary-weight variants. Exploratory software-only studies of HLF feature-token models based on pooling, MLP-mixing, and self-attention were also performed; these are not particle-level reproductions of the corresponding literature architectures.

This study provides a reproducible evaluation setup based on an existing public dataset and established FastML tools. The trained dense models, conversion artifacts, fixed evaluation data, and software reference predictions have been prepared for hls4ml and Vivado evaluation. Initial FPGA synthesis will target dense MLP models for which direct or explicit conversion routes have been prepared. The hardware study will evaluate resource utilization, synthesized latency, initiation interval, and estimated kernel throughput. It will determine whether the predictive-performance losses associated with aggressive weight quantization are offset by FPGA implementation benefits.

Do you plan to submit a 4-page extended abstract on OpenReview (only for Presentations/Posters)? Maybe

Author

Dorian Sloot (Austrian Academy of Sciences (AT))

Presentation materials

Peer reviewing

Paper