No news.
Dev machine: running at P2. Currently bare bones setup with CUDA and ROCm.
Slurm integration: Giada working on it, some compatibility issues because of mixing Alma 9 & 10 (delayed because of network issues at P2)
Managed to compile O2 and standalone benchmark.
Finished vectorizing kernel (GPU threads load 4-8 charges at once).
Performance barely changes...
Also old kernel was assuming wrong memory layout (pad-major instead of time-major).
Resulted in warps accessing two far apart cachelines. Fixing this also didn't change performance...
Next steps: Try micro benchmarks how to improve access pattern?