Speaker
Description
Custom ASIC accelerators offer significant power and performance advantages for machine learning in scientific and edge computing; a driving example is superconducting qubit readout, where moving real-time classification of qubit states from room-temperature FPGAs into the cryostat requires custom ASICs on cryo-compatible technology nodes. However, obtaining accurate area and timing requires synthesis runs that can take hours per design. This not only makes exhaustive design-space exploration infeasible but also limits emerging AI-driven and agentic design flows, in which an optimization loop may need to evaluate thousands of candidate designs. A surrogate model sidesteps the synthesis bottleneck by learning to predict synthesis metrics directly from a design specification and synthesis directives, without running synthesis at all. Such models were recently introduced for hls4ml accelerators on FPGA targets, but two challenges remain: extending the approach to ASIC synthesis and generalizing predictions to architectures outside the training distribution. We address both. We present a large-scale ASIC synthesis dataset and surrogate models that predict area, latency, and throughput for dense neural network accelerators generated by hls4ml and synthesized with Siemens Catapult HLS. The dataset contains over half a million designs targeting Nangate 45nm, spanning a wide range of network depths, layer widths, bitwidths, and reuse factors, on which we train and compare Transformer and Graph Neural Network surrogates. On held-out designs, our best model achieves $R^2 = 0.9997$ for latency and $R^2 = 0.9991$ for area; throughput is computed exactly rather than learned. These results, however, reflect interpolation within the training distribution. The second challenge surfaces under depth extrapolation: when trained on $n$-layer designs and tested on deeper $(n+m)$-layer networks, the same models degrade sharply; at $n=2$ and $m=1$, for example, latency $R^2$ falls to $0.634$ and area $R^2$ to $0.789$. The degradation suggests that, although the models interpolate accurately, they do not fully capture the structural overheads that each additional layer introduces, including control logic, pipeline coordination, and intermediate buffering. Integrating a sum-decomposition strategy into the model restores near-interpolation accuracy, with latency $R^2$ of $0.9954$ and area $R^2$ of $0.9797$. Finally, we demonstrate cross-node transfer: a model pre-trained with Nangate 45nm fine-tunes to the room-temperature GlobalFoundries 22FDX PDK using only $10\%$ of the pre-training data volume (${\approx}50{,}000$ designs), reaching latency $R^2$ of $0.9984$ and area $R^2$ to $0.9393$. This establishes open-library pre-training as a scalable data strategy for nodes whose PDKs are not yet ready for large synthesis campaigns, such as the cryogenic extension of 22FDX targeted for in-cryostat readout, and reduces the cost of evaluating a candidate design from hours of synthesis to milliseconds of inference, fast enough to embed in automated design-space exploration and agentic optimization loops.
| Do you plan to submit a 4-page extended abstract on OpenReview (only for Presentations/Posters)? | Yes |
|---|