Speaker
Description
High-level synthesis (HLS) has greatly improved the accessibility of FPGAs by enabling a faster transition from algorithmic descriptions to efficient hardware implementations. Advances in automated design space exploration (DSE) and MLIR-based compiler flows, such as ScaleHLS, have further enhanced the ability to transform high-level algorithms into optimized hardware designs. Recent research such as HIDA has extended these capabilities by automating the conversion of machine learning model graph structures into scalable dataflow architectures using HLS streams, demonstrating notable gains in both scalability and performance.
Despite this progress, current dataflow designs still depend on emulating data movement via configurable connections within the FPGA fabric. Consequently, when dataflow architectures are constructed to closely resemble the graph structure of machine learning models, they often encounter severe routing congestion as numerous nodes compete for limited connectivity resources. To mitigate this, the dataflow tool must substantially transform the original dataflow representation to fit the constraints of the more traditional von Neumann computing model. However, this transformation process introduces bottlenecks that ultimately limit both the performance and scalability of the resulting hardware designs.
To overcome these limitations, we introduce StreamFlex, a new framework that fully leverages the advanced Network-on-Chip (NoC) capabilities of the AMD Versal V80 FPGA. Rather than hardwiring dataflow designs onto the FPGA fabric, StreamFlex uses the native NoC to efficiently route data between hardware nodes, significantly reducing the need for aggressive abstraction-level transformations. Additionally, StreamFlex utilizes dynamic function exchange (DFX) to enable runtime reconfiguration of hardware blocks based on the active regions of the dataflow graph, maximizing hardware utilization and adaptability. This approach aims to closes the gap between flexibility and performance, making FPGA acceleration practical for a wide range of applications, including machine learning, cryptography, and high-performance computing, while lowering barriers to adoption for developers.