Speaker
Description
In many domains of science the likelihood function is a fundamental ingredient used to statistically infer model parameters from data, due to the likelihood ratio (LR) as an optimal test statistic. Neural based LR estimation using probabilistic classification has therefore had a significant impact in these domains, providing a scalable method for determining an intractable LR from simulated datasets via the so-called ratio trick [1,2].
The underlying paradigm of probabilistic machine learning adheres to the standard Kolmogorov axioms of probability theory [3], which requires the probability of an event in a measurable space to be nonnegative; a requirement met by classical systems. In contrast, quantum mechanical systems can be represented by quasiprobabilistic distributions, which allow for events with negative probabilities [4]. In high energy physics this is a significant problem when simulating proton-proton (pp) collisions using quantum field theory, due to the fact that Monte Carlo simulation codes can introduce negatively weighted data [5,6].
When using the aforementioned neural based ratio trick with negatively weighted data two problems present themselves. First, the variance of the mini-batch losses used during neural network parameter updates are systematically increased, thereby hindering the convergence of stochastic gradient descent (SGD) algorithms. The second is that most classification and density (ratio) estimation loss functions constrain the neural LR estimates to be in the range $[0,\infty)$. Therefore, should negative densities prevail anywhere within the measurable space, the neural network would be incapable of expressing such behaviour.
This work will demonstrate two important advancements for LR estimation with negatively weighted data. First, a new loss function for binary classification is introduced to extend the neural based LR trick to be compatible with quasiprobabilistic distributions. Second, signed probability spaces are used to decompose the likelihoods into signed mixture models. This decomposition reduces the overall LR estimation task into four nonnegative LR estimation sub-tasks, each with reduced loss variance during optimization relative to the overall task. Each nonnegative LR is estimated using a calibrated neural discriminative classifier [2], which are then combined via coefficients that are optionally optimised using the new loss function. The technique is demonstrated using di-Higgs production via gluon-gluon fusion in pp collisions at the Large Hadron Collider.
References
[1] Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori. Density Ratio Estimation in Machine Learning. Cambridge University Press, 2012.
[2] Kyle Cranmer, Juan Pavez, and Gilles Louppe. Approximating likelihood ratios with calibrated discriminative classifiers, 2016.
[3] A.N. Kolmogorov. Grundbegriffe der Wahrscheinlichkeitsrechnung. Number 1. Springer Berlin, Heidelberg, 1933.
[4] Richard Phillips Feynman. Negative probability. 1984.
[5] Stefano Frixione and Bryan R Webber. Matching nlo qcd computations and parton shower simulations. Journal of High Energy Physics, 2002(06):029–029, June 2002.
[6] Paolo Nason and Giovanni Ridolfi. A positive-weight next-to-leading-order monte carlo for Z pair hadroproduction. Journal of High Energy Physics, 2006(08):077–077, August 2006.
Significance
This abstract represents an original piece of work submitted to arXiv, with an on-going submission to MLST. It has not been presented within this community, and is not a status update/report of an on-going project.
References
arXiv: https://arxiv.org/abs/2410.10216 (MLST submission pending)
| Experiment context, if any | LHC Collider Physics |
|---|