Speaker
Description
Over the last few years, different pre-training strategies for foundation models in HEP have been proposed. Some of them, like generative pre-training (used in OmniJet-$\alpha$) and Masked Particle Modeling (MPM), rely on self-supervised pre-training, allowing models to be pre-trained on unlabelled data collected by experiments.
We present studies that compare those two self-supervised methods with straightforward supervised pre-training for jet tagging. Furthermore, we investigate how combinations of those different methods affect the quality of the pre-trained representations.
While the original OmniJet-$\alpha$ model used exclusively tokenized input, leading to a loss of information that could decrease the performance on downstream tasks such as jet tagging, we here use a hybrid setup with continuous input features for all tasks and tokenized particle representations as next-token-prediction target (OmniJet-$\alpha$) or masked-token-prediction target (MPM), respectively.