HSF DAAA: Agentic AI in HEP data analysis

Europe/Zurich
Jamie Gooding (Technische Universitaet Dortmund (DE)), Luke Kreczko (University of Bristol (GB)), Matthew Feickert (University of Wisconsin (US)), Nick Smith (Fermi National Accelerator Lab. (US))
Description

The HSF Data Analysis Working Group (DAAA) will be holding a meeting on the topic of agentic AI in data analysis. The meeting will consist of a short series of invited talks, followed by a freeform discussion on the arising topics. With this meeting we aim to understand the state of play in this exciting area and begin to establish a set of best practices for analysts who wish to incorporate these technologies in their own work.

Zoom Meeting ID
61591101367
Host
Jamie Gooding
Alternative hosts
Luke Kreczko, Nick Smith
Useful links
Join via phone
Zoom URL
    • 16:00 16:05
      Introduction 5m
      Speakers: Jamie Gooding (Technische Universitaet Dortmund (DE)), Dr Luke Kreczko (University of Bristol (GB)), Matthew Feickert (University of Wisconsin (US)), Nick Smith (Fermi National Accelerator Lab. (US))
    • 16:05 17:05
      Talks
      • 16:05
        AI Agents Can Already Autonomously Perform Experimental High Energy Physics 15m

        I will discuss our recent work (arXiv:2603.20179) showing that LLM-based agents, given a dataset, an execution framework, and prior literature, can autonomously carry out a complete HEP analysis, from event selection through statistical inference to paper drafting, demonstrated on ALEPH, DELPHI, and CMS open data with the Just Furnish Context (JFC) framework. I will then outline the way forward: rigorous benchmarking of agents against known physics results, and what routine agentic analysis would mean for how we work.

        Speaker: Andrzej Novak (Massachusetts Inst. of Technology (US))
      • 16:20
        Questions 5m
      • 16:25
        Osprey: Deploying Agentic AI Across Accelerator Facilities 15m

        Osprey is an agentic AI framework for operating particle accelerators. It turns a natural-language request into a plan the system carries out against the live control system: resolving which signals and devices the operator means, retrieving archived data, generating and running the necessary analysis and control code, and returning results; all under operator approval and bounded safety constraints. Osprey has been deployed across a range of DOE accelerator facilities, where it performs multi-stage machine-physics tasks that previously took an expert hours of scripting. This talk describes how the framework works, what it does on shift, how its safety constraints are enforced and tested, and what a shared, portable platform changes for bringing agentic AI into the control room.

        Speaker: Thorsten Hellert
      • 16:40
        Questions 5m
      • 16:45
        Benchmarking AI Agents for High-Energy Physics Analysis 15m

        This talk presents an evolving program to benchmark large language models and agentic systems on realistic high-energy physics analysis tasks. It begins with a 2025 study using a simple supervisor–coder framework to compare model performance across stages of a Higgs diphoton analysis, where the limited harness made differences largely attributable to the underlying models. The focus then shifts to modern general-purpose coding agents, evaluated as complete systems combining a model, harness, context, tools, and retry policies. Using Terminal-Bench and Harbor, we construct collider-analysis benchmarks covering top-associated Higgs production and hadronic top reconstruction. Evaluation is based on detailed physicist-defined questions about correctness, methodological compliance, validation practice, and scientific performance, with evidence extracted from run artifacts and scored deterministically. The resulting framework supports transparent comparisons of agent quality, reliability, cost, and required supervision, while also producing trajectories suitable for future supervised fine-tuning and reinforcement learning.

        Speakers: Haichen Wang (Lawrence Berkeley National Lab. (US)), Haichen Wang (UC Berkeley)
      • 17:00
        Questions 5m
    • 17:05 17:30
      Freeform discussion 25m