Madgraph5 GPU development - !! ATTENTION !! STARTS 16:00 !!

Europe/Zurich
513/R-068 (CERN)

513/R-068

CERN

19
Show room on map
Zoom Meeting ID
68039306359
Host
Stefan Roiser
Alternative hosts
Daniele Massaro, Kenneth Rioja, Frantisek Stloukal, Yann Zurstrassen, Enrico Bothmann, Olivier Mattelaer, Alexia Marie Yiannouli, Maksymilian Graczyk
Passcode
52022438
Useful links
Join via phone
Zoom URL

Meeting minutes 2026-09-01

Quick recap

This meeting focused on discussing performance tests and development updates for MadGraph 7.

  • Daniele presented results from recent performance tests comparing MadGraph 5 and MadGraph 7, showing generally stable cross-section ratios with small errors, though there were issues with the triple W process that appeared to be related to integration errors in MadGraph 5 rather than MadGraph 7 itself.
  • The team discussed plans for additional testing including running more processes with MadNIS training parameter train_batches set to 200K, and testing on different CPU architectures (also supporting AVX-512 with native 512 b registers) to better understand performance variations.
  • Theo reported progress on reproducibility improvements, fixing GPU performance issues related to random number generation that had been causing significant slowdowns, and noted that convergence issues in MadNIS training needed further investigation.
  • The group also discussed code splitting strategies, agreeing to proceed with implementing the renaming changes and then focusing on the merging of the code splitting before the alpha release
  • There was a discussion about the systematics module performances that was addressed in a separate discussion.

Next steps

Daniele

  • Run 5 MadGraph 7 processes with MadNIS train_batches set to 200K, including all missing processes that failed before, and adding V + jets (W+- jets), and repeat MadGraph 5 runs for comparison.
    Especially triple W process might behave better with sde_strategy 2, and/or FD gauge.
  • Rerun CPU tests on a machine with AVX512 and 512b registers support.
  • Redo plots showing MadGraph 5 and MadGraph 7 errors separately, and consider different plot representations.
  • Collect more runs with different seeds for MadGraph 7 to compare runtime variation with MadGraph 5 (probably not needed, we have already 10 measurements)
  • Coordinate the big renaming PR merging.

Olivier

  • Finish the MadSpin paper for MadGraph 5 and then work on the MadGraph 7 part.

Theo

  • Finalize and send the missing sections (MadSpace, and Madboard) for the paper.
  • Open a PR for the reproducibility feature once the MadNIS training convergence issue is fixed.
  • Aim to have final performance plots by mid-September for the ML4Jets plenary.

Collaboration

  • Daniele & Team: Close the two open PRs related to code splitting, then merge the code splitting PR.
  • Daniele & Olivier & Theo: Have a private chat to discuss the default setting for systematics (scale and PDF variations).

Summary

MadGraph Process Baseline Results Review

Daniele presented baseline results for MadGraph 5 and MadGraph 7 processes run on a machine with pinned CPUs, showing cross-section ratios and throughput data.
The team identified an issue with the 3W process in MadGraph 5, which had a larger error bar compared to other processes, likely due to problems with the MadGraph 5: better results could be achieved using sde_strategy 2 and/or FD gauge.
Daniele agreed to run additional tests with 200K batch size, include missing processes like V + jets, and rerun CPU tests on a machine with AVX-512 and 512b registers support, while also revising the plot presentation to show errors separately for MadGraph 5 and 7.

Tests Configuration and runtime variations

MadGraph 5 baselines were run with run.sh -p 8, and explained that collecting all the data takes about a week due to bottlenecks.
The team agreed that 10 data points were representative and not needed more, with a suggestion to compare results side by side with MadGraph 7.

Document and Performance Progress Update

The team discussed progress on a document and paper, with Daniele noting that metrics are complete but more details could be added if needed.
Theo reported on reproducibility work, identifying and fixing a performance issue with GPU runs where random number generation was significantly slower due to reseeding, which has now been optimized to run 100 times faster.
Theo also mentioned discovering a convergence bug in the MadNIS training that still needs investigation.

Alpha Release Planning Discussion

The team discussed plans for an upcoming alpha release, with Theo noting the need to have performance results by mid-September for an upcoming presentation.
Olivier reported progress on MadGraph 5 and MadGraph 7 releases, including work on the MadSpin paper.
The team agreed to prioritize closing two existing PRs blocking the code splitting PR before merging new changes, with the PR to be implemented as soon as possible rather than waiting for the alpha release.

GPU Performance Optimization Strategies

The team discussed performance optimization strategies for GPU implementation, including splitting aloha objects into separate momenta and wave functions as suggested by Olivier.
Frank reported that storing phase space points in shared memory and computing different helicity combinations within the same block by different threads did not provide performance gains.
Theo identified an issue with the current 2D grid implementation where empty blocks can occur due to cuts, suggesting a potential improvement by flattening the grid into a single dimension to avoid multiple partially filled blocks.

Frank also explained his findings on L2 cache allocation and color summation within the same kernel block, while the group explored the feasibility of diagonalizing color metrics to simplify calculations.

The discussion also covered Monte Carlo event generation costs, where Stefan noted a significant discrepancy between his estimates (hundreds of billions) and CMS's actual requirements (125 billion events), with Olivier suggesting that the bulk of events might not be as computationally intensive as expected.

CPU Device Naming Strategy Update

The team discussed and agreed on a new approach for CPU device naming, deciding to have separate variables for CPU mode rather than a single device option.
The team debated whether to keep systematic scale variations enabled by default, with some arguing it's important for physics despite being slow, while others suggested temporarily disabling it until hardware acceleration (thanks to Zenny's work) is available.

Reports for loop calculations

They also reviewed GPU performance improvements for the oneloop calculations, noting that while current speedups are lower than previous versions due to kernel execution issues, the implementation now runs the full calculation successfully.

There are minutes attached to this event. Show them.
    • 16:00 16:10
      News 10m
    • 16:10 16:30
      Topical discussion 20m
      Speaker: Dr All
    • 16:30 16:50
      Round table 20m
    • 16:50 17:00
      AoB 10m