Speaker
Description
Multi-latent attention (MLA) shrinks the KV cache but must rebuild per-head keys and values during decode, shifting the bottleneck to on-chip movement and orchestration. Spatial architectures, which consist of many-core tiles with local memory, explicit movers, and on-chip networks, therefore reward careful dataflow and mapping as much as computing. Unlike FPGA-centric stacks with long place-and-route time, the spatial architecture platforms support faster map–measure–refine loops.
We summarize two lines of work. On Tenstorrent Wormhole architecture, we treat MLA decode as a coupled mapping and scheduling problem under decoupled read–compute–write execution. Methodologically, we use a parameterized dataflow template, a cost model calibrated with simple empirical corrections, and an auto-tuner to explore a large, regime-dependent design space systematically. Measured decode behavior changes character with context length—from compute-dominated at short sequences to a movement- and coupling-limited regime at long context—motivating architecture-aware redesign rather than incremental tuning of a single fixed mapping.
On AMD Ryzen AI (Strix Halo), we build end-to-end MLA decode through a constraint-driven co-design loop: memory layout, DMA tiling, and matrix-multiply kernels are co-optimized so the full decode step runs as a single coordinated invocation across multiple compute columns, avoiding fragile multi-stage glue that often dominates latency on tightly coupled spatial substrates.
As physics-driven AI moves toward larger contexts, tighter latency budgets, and on-detector or facility-adjacent inference, spatial architectures offer a concrete path to turn memory-efficient attention into sustained throughput. https://ceca.pku.edu.cn/en/people_/faculty_/guojie_luo/