Speaker
Description
The search for new physics at the Large Hadron Collider (LHC) increasingly depends on the ability to extract tiny signals from petabytes of data. Machine learning (ML) methods have become essential tools for identifying jets and the particles that initiated them, as well as for separating rare physics processes from background events.
In recent years, self-supervised learning (SSL) has been explored for jet classification, particularly due to its ability to leverage large amounts of unlabeled data. This approach is appealing because on model can replace training many single-task models with supervised learning.
In this work, we introduce JP-JEPA, a novel framework based on joint embedding predictive architectures (JEPA). Our approach aims to predict a latent representation associated with an enriched view of the data from a partial observation. Built upon the state-of-the-art Particle Transformer (ParT), our method captures fine-grained correlations among jet constituents without requiring explicit supervision.
Using the JetClass open dataset, we demonstrate state-of-the-art performance compared to supervised methods. Additionally, we propose and validated a feature-level masking strategy tailored to missing information and other experimental effects. Our results demonstrate that a competitive pre-trained foundation model can be built for downstream applications in Particle Physics.
| I read the instructions above | Yes |
|---|