KaleidoForge Logo

KalEdge-Lite: Hardware-Aware ML-to-FPGA Deployment with Automated hls4ml Integration

Bridging Neural Network Compression and FPGA Deployment in a Unified Workflow

Romina Soledad Molina

KaleidoForge, Italy

hls4mlVitis HLSTF-MOTQKerasVivado

The Integration Problem

"The barrier is not missing algorithms, it is missing integration."

End-to-End Workflow

1
Load & Analyze
Upload dataset, AI recommends preprocessing
2
Design Architecture
Define model via GUI or AI Architect
3
Train Baseline
Float32 training with K-Fold CV
4
Optimize & Compress
Pruning, KD, QAT, SVD, or fused (KD+QAP) pipelines
5
Evaluate
Accuracy vs. Size vs. Params comparison
6
Convert to HLS
Keras to C++ via hls4ml, resource estimation
7
Deploy
DMA wrappers + Build Agent synthesis

Validated with ~70 users across FPGA and ML courses in Argentina, Italy, Spain, and online

All config auto logged as JSON. Deterministic regeneration of trained models, HLS projects, and bitstreams

Local tunnel invokes Vitis HLS and Vivado on the user's machine. No cloud licensing required

KalEdge Platform

Web-based dashboard orchestrating the full ML to FPGA pipeline

KalEdge Platform

Compression Suite

Technique Engine FPGA Impact Combinable
Baseline Training Keras / K-Fold CV Accuracy ceiling reference
Pruning TF-MOT Fewer LUTs / FFs via sparsity Yes
QAT TF-MOT / QKeras hls4ml ready fixed-point Yes
Knowledge Distillation Custom Accuracy recovery in small models Yes
Low-Rank (SVD) Custom Fewer parameters, fewer DSPs Yes
KD + QAP Fused pipeline Max compression, single run

Scoring profiles, configurable weights for Accuracy / Size / Params:

FPGA-oriented Embedded Ultra-TinyML Accuracy-driven

Regime-Aware Analytical Surrogate

Predicts LUT, FF, DSP, and latency without invoking synthesis

Compute-boundR ≤ P → Latency ≈ const
Reuse-boundR > P → Latency scales with R
Latency-dominatedPipeline fill overhead

Analytical estimation

Models how FPGA resources scale with model size, precision, and parallelism.

Board-family aware

Separate coefficient sets for Zynq-7020 (DSP48E1) vs. UltraScale+ (DSP48E2), auto selected from UI

Infeasibility detection

Over-budget configurations flagged before synthesis, enables systematic DSE rejection

Estimation Equations

Mathematical primitives mapping topology and reuse factor to hardware cost

DSP Block Utilization

DSP usage scales inversely with reuse factor R, adjusted by an efficiency factor κ(R):

DSP(M, R) = ⌈κ(R) · M / R⌉
Where κ(R) = 0.72 + 0.28 · (1 - e-R/16). This empirical correction factor (fitted from synthesis data) accounts for routing overhead at low R, reaching ideal scaling as R increases.

Logic & Latency Models

Logic resources scale linearly with MAC count, while latency bridges regimes analytically:

LOGIC (LUT & FF)
LUT(M) = β₀ + αlut · M FF(M) = (γ₀ + αff · M) · uplift
LATENCY (CYCLE COUNT) Lat(M, R) = (a + b·R) · (ΣSout/Sref)α · corr
Latency includes corrections for pipeline depth and pipeline-fill overhead for small networks.
M: total MACs (weights)R: Reuse Factorκ(R): routing efficiencyβ0 / αlut / γ0 / αff: logic parametersuplift: FF correction factor
a, b, α: latency coefficientsΣSout / Sref: total output dimension ratiocorr: small-network correction factor
For this presentation, only the following hls4ml configurations are shown: ReuseFactor, io_parallel, and Latency as strategy. Quantization: 16/16 and 8/16. The tool also supports io_stream and Resource as strategy.
Work in progress, new hls4ml features are being progressively added to the surrogate.

Design-Space Exploration Speedup

2–20 min
HLS Synthesis per config
<10 ms
KalEdge surrogate per config
103–104x
Speedup over synthesis
Method Time / Config Total (20 configs)
HLS Synthesis 2–20 min 40–400 min
KalEdge Surrogate < 10 ms < 200 ms

Surrogate Accuracy

Mean Absolute Percentage Error (MAPE) of KalEdge Estimator relative to Vitis HLS 2024.1 across R ∈ {1, 8, 16, 32} (all sub-10% after calibration)

Workload Device W/A DSP % FF % LUT % Lat. %
G/N MLP (536p) Zynq-7020 16/16 0.0 0.9 2.6 3.2
G/N MLP (536p) Zynq-7020 8/16 0.0 3.8 2.7 1.6
MNIST MLP (3270p) * Zynq-7020 8/16 0.0 7.7 4.0 0.4
MNIST MLP (6866p) * Zynq-7020 8/16 1.5 1.0 0.8 2.6
G/N MLP (536p) UltraScale+ 16/16 0.0 1.1 2.2 2.0
G/N MLP (536p) UltraScale+ 8/16 0.0 6.0 2.0 4.4
MNIST MLP (3270p) * UltraScale+ 8/16 0.0 3.7 1.0 1.3
MNIST MLP (6866p) * UltraScale+ 8/16 1.1 1.2 0.8 1.1

* Exceeds board capacity; surrogate correctly classifies as infeasible during DSE (Note: downstream Vivado logic optimizations may still recover feasibility).

Zynq-7020 HLS vs. KalEdge Estimates

Vitis HLS 2024.1 vs. KalEdge Surrogate on Zynq-7020

Model W/A RF DSP FF LUT Latency
HLS Est. HLS Est. HLS Est. HLS Est.
G/N (536p) 16/16 1 422 422 16,997 16,997 15,814 15,814 22 22
G/N (536p) 16/16 32 20 20 11,696 11,862 17,250 17,205 72 72
G/N (536p) 8/16 1 9 9 21,461 21,461 13,787 13,787 46 46
G/N (536p) 8/16 32 4 4 7,919 8,342 8,386 8,555 49 49
MNIST (3270p) 8/16 1 25 25 135,647 135,647 92,482 92,482 139 138
MNIST (3270p) 8/16 32 25 25 59,592 67,593 56,139 61,423 154 154
MNIST (6866p) 8/16 1 1,598 1,598 110,075 110,000 242,914 244,000 25 24
MNIST (6866p) 8/16 32 205 202 99,827 98,375 239,623 240,125 92 91

Exceeds board resource limits (220 DSPs, 106k FFs, 53.2k LUTs); surrogate correctly flags as infeasible.

UltraScale+ HLS vs. KalEdge Estimates

Vitis HLS 2024.1 vs. KalEdge Surrogate on xczu3eg

Model W/A RF DSP FF LUT Latency
HLS Est. HLS Est. HLS Est. HLS Est.
G/N (536p) 16/16 1 397 397 7,264 7,264 15,502 15,502 9 9
G/N (536p) 16/16 32 20 20 11,259 11,066 17,638 17,688 64 69
G/N (536p) 8/16 1 10 10 4,107 4,292 15,741 15,855 9 9
G/N (536p) 8/16 32 5 5 4,425 4,100 15,906 15,500 14 14
MNIST (3270p) 8/16 1 25 25 86,388 86,388 85,669 85,669 35 34
MNIST (3270p) 8/16 32 25 25 25,034 25,868 57,217 57,799 52 52
MNIST (6866p) 8/16 1 1,598 1,596 33,291 33,291 245,738 247,600 14 14
MNIST (6866p) 8/16 32 205 199 54,599 53,266 244,398 247,600 84 86

Exceeds board resource limits (360 DSPs, 141k FFs, 70.5k LUTs); surrogate correctly flags as infeasible.

Discussion & Limitations

What works now

MLP surrogate: sub 10% MAPE on Zynq-7020 and UltraScale+. Dashboard tested and validated with ~70 participants.

Three AI agents

Dataset Agent, AI Architect, AI Advisor: abstract preprocessing, architecture scripting, and hardware config decisions.

Current scope

Surrogate calibrated for MLP only. CNN extension planned. LRF and sparsity (pruning) impacts not yet integrated into hardware estimation.

Future work

CNN surrogate, LRF hardware coupling, sparsity-aware resource estimation, hardware guided training loop, SNN (LIF) HLS pipeline. Recalibrate surrogate against Vivado post-implementation results for final resource accuracy.

Conclusions

Live Demo

Launch KalEdge Platform

Thank You

Questions & Discussion

QR Contact

Romina Soledad Molina

rsmolina@kaleidoforge.com

www.kaleidoforge.com

QR KalEdge

KalEdge Platform

kaledge.kaleidoforge.com

Automated ML-to-FPGA Optimization & Deployment

1 / 15