Speaker
Description
Abstract
Deploying neural networks on FPGAs remains a significant barrier in embedded AI. Exploring the design space across compression strategies (pruning, quantization, knowledge distillation) and hardware configurations is often fragmented, requiring deep expertise across multiple toolchains.
This tutorial introduces KalEdge, a hardware-aware platform unifying the ML-to-FPGA workflow. Attendees will orchestrate architecture definition, compression-aware training, integration with hls4ml, and bitstream generation from a single interface. Through hands-on exercises, participants will compress and deploy a neural network onto a physical FPGA (HyperFPGA cluster).
Detailed Outline
Part 1: Introduction and Theory (20m)
• The ML-to-FPGA deployment gap and KalEdge platform architecture.
• Hardware-aware compression basics and the analytical surrogate model.
Part 2: App Tutorial & Basic Exercises (30m)
• Introduction to the KalEdge UI.
• Hands-on with basic guided exercises to familiarize users with the platform's end-to-end workflow.
Part 3: Advanced Pipeline & Exploration (45m)
• Loading datasets, defining architectures, and executing the automated pipeline.
• Rapid Design Space Exploration (DSE) using the surrogate estimator.
Coffee Break (15m)
Part 4: HLS Generation and Deployment (60m)
• Converting models to C++ via hls4ml and generating firmware wrappers (AXI-Stream DMA).
• Local Synthesis Tunnel: Generating bitstreams via Vivado/Vitis HLS (pre-synthesized IPs and XSA files will be provided for participants without local installations).
• Hardware-in-the-Loop: Remote deployment and live inference visualization using the HyperFPGA cluster (MLab-STI-ICTP).
Part 5: Conclusion & Q&A (10m)
• Summary and future directions.
Infrastructure & Validation
• Required: Attendees need a laptop with a web browser and SSH client.
• Provided: Cloud-hosted backend instances, pre-synthesized IPs and XSA files (for users without local Vivado installations), access to the HyperFPGA cluster (MLab-STI-ICTP) for remote deployment, and a pre-configured Virtual Machine (VM) if necessary.
Duration: 3hs
Target Audience: Researchers, students, and practitioners in Edge AI and hardware acceleration. Basic Python and ML knowledge required; no RTL/HLS expertise needed.
| Tutorial level (only for Tutorial) | Intermediate |
|---|