OpenCL

No news.

GPU Servers

Dev machine: running at P2. Currently bare bones setup with CUDA and ROCm.

Slurm integration: Giada working on it, some compatibility issues because of mixing Alma 9 & 10 (delayed because of network issues at P2)

Managed to compile O2 and standalone benchmark.

Highly Ionizing Particles

Finished vectorizing kernel (GPU threads load 4-8 charges at once).

Performance barely changes...

Also old kernel was assuming wrong memory layout (pad-major instead of time-major).

Resulted in warps accessing two far apart cachelines. Fixing this also didn't change performance...

Next steps: Try micro benchmarks how to improve access pattern?