Bridging Neural Network Compression and FPGA Deployment in a Unified Workflow
Romina Soledad Molina
KaleidoForge, Italy
Validated with ~70 users across FPGA and ML courses in Argentina, Italy, Spain, and online
All config auto logged as JSON. Deterministic regeneration of trained models, HLS projects, and bitstreams
Local tunnel invokes Vitis HLS and Vivado on the user's machine. No cloud licensing required
Web-based dashboard orchestrating the full ML to FPGA pipeline
| Technique | Engine | FPGA Impact | Combinable |
|---|---|---|---|
| Baseline Training | Keras / K-Fold CV | Accuracy ceiling reference | |
| Pruning | TF-MOT | Fewer LUTs / FFs via sparsity | Yes |
| QAT | TF-MOT / QKeras | hls4ml ready fixed-point | Yes |
| Knowledge Distillation | Custom | Accuracy recovery in small models | Yes |
| Low-Rank (SVD) | Custom | Fewer parameters, fewer DSPs | Yes |
| KD + QAP | Fused pipeline | Max compression, single run |
Scoring profiles, configurable weights for Accuracy / Size / Params:
FPGA-oriented Embedded Ultra-TinyML Accuracy-drivenPredicts LUT, FF, DSP, and latency without invoking synthesis
Models how FPGA resources scale with model size, precision, and parallelism.
Separate coefficient sets for Zynq-7020 (DSP48E1) vs. UltraScale+ (DSP48E2), auto selected from UI
Over-budget configurations flagged before synthesis, enables systematic DSE rejection
Mathematical primitives mapping topology and reuse factor to hardware cost
DSP usage scales inversely with reuse factor R, adjusted by an efficiency factor κ(R):
DSP(M, R) = ⌈κ(R) · M / R⌉
κ(R) = 0.72 + 0.28 · (1 - e-R/16). This
empirical correction factor (fitted from synthesis data) accounts for routing overhead at low R, reaching
ideal scaling as R increases.
Logic resources scale linearly with MAC count, while latency bridges regimes analytically:
LUT(M) = β₀ + αlut · M
FF(M) = (γ₀ + αff · M) · uplift
Lat(M, R) = (a + b·R) · (ΣSout/Sref)α · corr
| Method | Time / Config | Total (20 configs) |
|---|---|---|
| HLS Synthesis | 2–20 min | 40–400 min |
| KalEdge Surrogate | < 10 ms | < 200 ms |
Mean Absolute Percentage Error (MAPE) of KalEdge Estimator relative to Vitis HLS 2024.1 across R ∈ {1, 8, 16, 32} (all sub-10% after calibration)
| Workload | Device | W/A | DSP % | FF % | LUT % | Lat. % |
|---|---|---|---|---|---|---|
| G/N MLP (536p) | Zynq-7020 | 16/16 | 0.0 | 0.9 | 2.6 | 3.2 |
| G/N MLP (536p) | Zynq-7020 | 8/16 | 0.0 | 3.8 | 2.7 | 1.6 |
| MNIST MLP (3270p) * | Zynq-7020 | 8/16 | 0.0 | 7.7 | 4.0 | 0.4 |
| MNIST MLP (6866p) * | Zynq-7020 | 8/16 | 1.5 | 1.0 | 0.8 | 2.6 |
| G/N MLP (536p) | UltraScale+ | 16/16 | 0.0 | 1.1 | 2.2 | 2.0 |
| G/N MLP (536p) | UltraScale+ | 8/16 | 0.0 | 6.0 | 2.0 | 4.4 |
| MNIST MLP (3270p) * | UltraScale+ | 8/16 | 0.0 | 3.7 | 1.0 | 1.3 |
| MNIST MLP (6866p) * | UltraScale+ | 8/16 | 1.1 | 1.2 | 0.8 | 1.1 |
* Exceeds board capacity; surrogate correctly classifies as infeasible during DSE (Note: downstream Vivado logic optimizations may still recover feasibility).
Vitis HLS 2024.1 vs. KalEdge Surrogate on Zynq-7020
| Model | W/A | RF | DSP | FF | LUT | Latency | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| HLS | Est. | HLS | Est. | HLS | Est. | HLS | Est. | |||
| G/N (536p) | 16/16 | 1 | 422 | 422 | 16,997 | 16,997 | 15,814 | 15,814 | 22 | 22 |
| G/N (536p) | 16/16 | 32 | 20 | 20 | 11,696 | 11,862 | 17,250 | 17,205 | 72 | 72 |
| G/N (536p) | 8/16 | 1 | 9 | 9 | 21,461 | 21,461 | 13,787 | 13,787 | 46 | 46 |
| G/N (536p) | 8/16 | 32 | 4 | 4 | 7,919 | 8,342 | 8,386 | 8,555 | 49 | 49 |
| MNIST (3270p) | 8/16 | 1 | 25 | 25 | 135,647† | 135,647† | 92,482† | 92,482† | 139 | 138 |
| MNIST (3270p) | 8/16 | 32 | 25 | 25 | 59,592 | 67,593 | 56,139 | 61,423 | 154 | 154 |
| MNIST (6866p) | 8/16 | 1 | 1,598† | 1,598† | 110,075† | 110,000† | 242,914† | 244,000† | 25 | 24 |
| MNIST (6866p) | 8/16 | 32 | 205 | 202 | 99,827 | 98,375 | 239,623† | 240,125† | 92 | 91 |
† Exceeds board resource limits (220 DSPs, 106k FFs, 53.2k LUTs); surrogate correctly flags as infeasible.
Vitis HLS 2024.1 vs. KalEdge Surrogate on xczu3eg
| Model | W/A | RF | DSP | FF | LUT | Latency | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| HLS | Est. | HLS | Est. | HLS | Est. | HLS | Est. | |||
| G/N (536p) | 16/16 | 1 | 397 | 397 | 7,264 | 7,264 | 15,502 | 15,502 | 9 | 9 |
| G/N (536p) | 16/16 | 32 | 20 | 20 | 11,259 | 11,066 | 17,638 | 17,688 | 64 | 69 |
| G/N (536p) | 8/16 | 1 | 10 | 10 | 4,107 | 4,292 | 15,741 | 15,855 | 9 | 9 |
| G/N (536p) | 8/16 | 32 | 5 | 5 | 4,425 | 4,100 | 15,906 | 15,500 | 14 | 14 |
| MNIST (3270p) | 8/16 | 1 | 25 | 25 | 86,388 | 86,388 | 85,669† | 85,669† | 35 | 34 |
| MNIST (3270p) | 8/16 | 32 | 25 | 25 | 25,034 | 25,868 | 57,217 | 57,799 | 52 | 52 |
| MNIST (6866p) | 8/16 | 1 | 1,598† | 1,596† | 33,291 | 33,291 | 245,738† | 247,600† | 14 | 14 |
| MNIST (6866p) | 8/16 | 32 | 205 | 199 | 54,599 | 53,266 | 244,398† | 247,600† | 84 | 86 |
† Exceeds board resource limits (360 DSPs, 141k FFs, 70.5k LUTs); surrogate correctly flags as infeasible.
What works now
MLP surrogate: sub 10% MAPE on Zynq-7020 and UltraScale+. Dashboard tested and validated with ~70 participants.
Three AI agents
Dataset Agent, AI Architect, AI Advisor: abstract preprocessing, architecture scripting, and hardware config decisions.
Current scope
Surrogate calibrated for MLP only. CNN extension planned. LRF and sparsity (pruning) impacts not yet integrated into hardware estimation.
Future work
CNN surrogate, LRF hardware coupling, sparsity-aware resource estimation, hardware guided training loop, SNN (LIF) HLS pipeline. Recalibrate surrogate against Vivado post-implementation results for final resource accuracy.
Questions & Discussion
Romina Soledad Molina
rsmolina@kaleidoforge.com
www.kaleidoforge.com