Alice Weekly Meeting: Software for Hardware Accelerators

→ Europe/Zurich
Zoom Meeting ID
61230224927
Host
David Rohr
Useful links
Join via phone
Zoom URL
    • 10:00 → 10:20
      Discussion 20m
      Speaker: David Rohr (CERN)

      Color code: (critical, news from this week: blue, news from last week: purple, no news: black)

      Sync reconstruction

       

      Async reconstruction

      • Need to investigate short GPU stall problem.
      • Limiting factor for pp workflow is now the TPC time series, which is to slow and creates backpressure (costs ~20% performance on EPNs). Enabled multi-threading as recommended by Matthias - need to check if it works.
      • Will tune existing 16-core settings, add a SITEARCH for 16core CPU, and 16coreCPU + generic NVIDIA / AMD GPU, like for 8 core.
      • Will retune EPN async workflow for TPC + ITS on GPU on 2025 data.

       

      GPU ROCm / compiler topics:

      • Problem with building ONNXRuntime with MigraphX support.
      • Need to find a way to build ONNXRuntime with support for CUDA and for ROCm.
      • Try to find a better solution for the problem with __device__ inline functions leaking symbols in the host code.
      • First version of ALIBUILD_O2_FORCE_GPU=build env variable option to use CUDA / ROCm from alidist merged.
        • Working nicely so far.
        • Out of the 2 bugs reported to ONNXRuntime for CUDA architectures, 1 is no fixed upstread, we can remove the workaround after the next update.
        • ALIBUILD_O2_FORCE_GPU=ci now defaults to "build" for SLC10. slc10-gpu-ci is now working in this mode, and seems good. (only thing still to be done: need to check that the architecture override is now fixed correctly after Giulio's last change).
        • Next to do: Move FullCI to SLC10 for O2 and alidist: Then we have all CI GPU builds on SLC10.
      • CUDA 13.4 released, supports GCC 16, will bump in cuda.sh recipe, but for now let's first stabilize the CI with the build mode.
      • Did a quick validation of ROCm10, seems as stable as 7.14, but still need to do long-term test.

       

      TPC / GPU Processing 

      • WIP: Use alignas() or find a better solution to fix alignment of monte carlo labels: https://its.cern.ch/jira/browse/O2-5314
      • Need to check the problem with ONNX external memory allocator.
      • Next high priority topic: Improvements for cluster sharing and cluster attachment at lower TPC pad rows. PR: https://github.com/AliceO2Group/AliceO2/pull/14542
      • Check for unnecessary f64 instructions in GPU code.

       

      Other topics:

      • GCC bump - Status?
      • We can go to GCC 16, once we switch the FullCI to SLC10 and the new GPU runtime build mode.

       

      EPN GPU Topics:

      • Discuss with Ernst / Giulio how to do a build for SLC10 on EPNs with build mode.
    • 10:20 → 10:25
      TPC ML Clustering 5m
      Speaker: Christian Sonnabend (CERN)

      alidist

      Successful merge of https://github.com/alisw/alidist/pull/6329 which now tests ONNXRuntime in the GPU CI with CUDA, MIGraphX and CPU execution providers

       

      momentum vector estimate for seeding

      • PR is still open: https://github.com/AliceO2Group/AliceO2/pull/15756
      • Plots for padrow0 show no difference (left is NN with mom. vector estimate, right is GPU CF. 0-100% centrality, Pb--Pb):
        • X: lowest pad row containing a native cluster with the matched MC track’s label
        • Y: reconstructed track’s innermost row, taken as the smaller row of its first and last stored hits

       

      -> No change there and similar when investigating against occupancy and also for 0-5% centrality enforced sim. Upshot: This was good to check, but NN momentum vector or current implementation really doesn't make a lot of difference.

       

      Brevitas quantization - MI300X (thanks to Oliver for the help)

      • Inference speedup

      -> dominated by the memory copies of ONNXRuntime

      -> Larger batch sizes can take advantage and then model calculation dominates

      • Model size scaling

      • accuracy vs. throughput

      • Memory footprint

      • Deployment cost

       

      Summary:

      Small models take less VRAM and have higher compute throughput. The validation QAT is much better than PTQ and can achieve similar quality for the cluster finding task (accuracy vs. throughput plot), but takes more training time (deployment cost)

       

       

    • 10:25 → 10:30
      GPU Parameter Optimizations 5m
      Speaker: Gabriele Cimador (CERN, Università and INFN Torino)
    • 10:30 → 10:35
      Efficient Data Structures 5m
      Speaker: Dr Oliver Gregor Rietmann (CERN)
       

      Implement NGT SoA Code in O2 standalone benchmark

      • Working on this fork of the AliceO2 repo, with a CI pipeline:
        • Running on NGT hardware with 4 different GPUs (Nvidia and AMD)
        • Extended CI-pipline to fail if GPU.out changes
      • Implemented SoA in SectorTracker: GPUTPCBaseTrackParam, GPUTPCTrackParam, GPUTPCTracklet, GPUTPCTrack

       

      • Make better use of SoA to improve performance
        • TrackletConstructor
          • Register spill from rowHits array, but the accesses are not coalesced. (UpdateTracklet)
          • global memory loads on average use only 4.4 out of 32 bytes sector
          • active warps per thread 9.34 out of 32 (29%) due to different tracklet lengths
          • Making the threads in a warp row-synchronous made performance only considerably worse.
        • ExtrapolationTracking
          • Register spill from rowHits array, but the accesses are not coalesced. (UpdateTracklet)
          • global memory loads on average use only 4.4 out of 32 bytes sector.
          • active warps per thread changed from 2.65 out of 32 (8%) to 
          • I made TrackletSelector filter the Tracklets that cross sector boundaries, so that we extrapolate only those.

       

      metric old new
      active warps per thread 3 out of 32 (8%) 23 out of 32 (73%)
      average SM occupancy 26% 10%
      bytes utilized per loaded sector 5 out of 32 12 out of 32
      register spill to lcoal memory (bytes) 1,493,529 19,549 (factor 76 better)

       

      • Next Steps
        • Fix coalescing by making the threads of a warp process the same rows
        • (Generate GPU parameters .csv file for SoA and to see if the optimal parameters are different)
    • 10:40 → 10:45
      TPC Clusterization / OpenCL / Highly Ionizing Particles 5m
      Speaker: Felix Weiglhofer (CERN)
    • 10:45 → 10:50
      ITS Tracking 5m
      Speakers: Felix Schlepper (CERN, Heidelberg University (DE)), Gabriele Cimador (CERN, Università and INFN Torino), Matteo Concas (CERN)
    • 10:50 → 10:55
      System Run Coordination Topics 5m
      Speaker: Ernst Hellbar (CERN)