Speaker
Description
Autoregressive generative models form the basis of modern LLMs. These models utilize a transformer architecture which requires caching the previous attention matrix (KV-Cache) for efficient inference. FPGAs are uniquely capable of high throughput due to their on-chip memory for both weights and biases as well as for KV-Cache. We present a GPT based autoregressive model trained using High Granularity Quantization (HGQ) on the TinyStories dataset and deployed through hls4ml on the Altera Agilex 7 accelerator card achieving ~100000 tokens per second. This work opens the avenue for FPGA based models for efficient time-series generation in fields where simulation consumes large amounts of computational resources including many branches of science.
| Do you plan to submit a 4-page extended abstract on OpenReview (only for Presentations/Posters)? | Maybe |
|---|