1–5 Sept 2025
ETH Zurich
Europe/Zurich timezone

Towards Fast and Interpretable Physics-Informed Learning: Second-Order Neurons and Mixed-Activation Networks

Not scheduled
20m
HIT G floor (gallery)

HIT G floor (gallery)

Speaker

João Paulo De Souza Böger

Description

Complex simulators are central to scientific research, forecasting, and real-world applications. However, they often require intensive computational resources and suffer from scalability issues — challenges amplified in the big data era. The APEX project addresses these limitations by designing novel efficient architectures for scientific simulators, exploring inductive biases, causal relations, and architectural innovations that learn and generalize simulator dynamics in a fast, scalable, uncertainty-aware, and interpretable way.

Our goal is to build surrogate models (or metamodels) that generalize beyond their training data, even when conditions change — a common challenge where standard machine learning (ML) models often fail. Many simulators, like those used in transport, climate modeling, epidemiology, and hydrodynamics, are governed by ordinary and partial differential equations (ODEs/PDEs). Efficiently learning these dynamics is crucial for effective surrogate modeling. To address the limitations of traditional ML approaches, we introduce inductive biases that guide learning based on known constraints of theses systems. While traditional PDE solvers and vanilla ML models struggle with scalability, determining which biases best help models generalize ODE/PDE dynamics remains an open challenge. This calls for model compression and reduced-parameter architectures that effectively blend expert knowledge with ML to create fast ML solutions for dynamical system simulators.

Quadratic neurons and domain-specific activation functions (MixFunn) [1], motivated by the solutions’ analytic forms to differential equations, provide a method for integrating domain insights into a neural network architecture. Coupled with soft constraints incorporated in the loss function as suggested by Physics-Informed Neural Networks (PINNs) [2], this innovative design matches or surpasses the performance of vanilla PINNs, yet utilizes considerably fewer parameters — facilitating closed-form solutions and enhancing model efficiency, which in turn can accelerate scientific discovery.

Building on this idea, we hypothesize that polynomial neurons and tailored non-linearities are key to creating scalable, generalizable architectures for scientific simulators. Our contributions include an extensive evaluation of MixFunn, along with PINNs, neuralODEs [3], and Equation Learners [4], in four benchmark ODE-based simulators. We use two recent proposed metrics [5] better suited for assessing learning quality OOD in surrogate models and test their robustness to noise in data. Furthermore, we introduce a test-time adaptation module [6], which allows models to refine parameters when limited trajectory data is available, improving accelerated inference for scientific applications.

We evaluate all approaches on widely used benchmarks: the SIR epidemic model, the Lotka–Volterra predator-prey system, Duffing oscillator, and Van der Pol oscillator. These systems span diverse tasks and dynamics, providing a robust testbed for comparing inductive biases, OOD generalization, and computational efficiency.

This work highlights the importance of architectural choices and domain-informed priors to closing the gap between robustness, speed, and generalization in scientific surrogates. It paves the way for compact, interpretable, and fast metamodels that encode scientific structure for real-world simulation tasks.

Author

Co-authors

Presentation materials