Perlmutter: Lots (~ thousands) of jobs failed due to SLURM job timeout -
- Especially during high parallel job periods (eg. 3600 pilots running on 230,400 cores all at once). Job finished, but the pilot is still running when SLURM time ends
- (Doug) This is expected behavior because pilot is processing work on jobs when the Slurm time limit is trigger. We get decent throughput with jobs running 6 hrs vs 24 hrs. we allow any Slurm jobs to run between a minimum of 6 hrs up to a max of 24 hrs
TACC: TACC is ok we ran the harvester on the Stampede3 login node. Need to reduce the number of threads used in the modules to keep the total under the user limit.
There is a flag MR (shared by Serhan) that can ask Athena to stop during event processing with a given type signal --> could be helpful in case of SLURM timeout in the middle of the execution stage. Testing it on Perlmutter.
- What is the flag? We have to see how the flag will be propogated to the pilot inside of a container