Speaker
Description
Transformers excel at modeling correlations in LHC collisions but incur high costs from quadratic attention. We analyze the Particle Transformer using attention maps and pair correlations on the (η,ϕ) plane, revealing that Particle Transformer attention maps learn traditional jet substructure observables. To improve efficiency we benchmark linear attention variants on JetClass and find that the Linformer matches ParT accuracy with far fewer resources. We then develop a novel small linear model informed by physics that partitions particles by kinematics and uses convolutional layers to capture local features. This model outperforms the standard Linformer on HLS4ML and has much lower latency than full attention models. Sequence ordering guided by physics and embedding analysis further improve accuracy and transparency for real time collider applications.