TACC
- Running jobs on the scale of ~10 nodes (1 node per slurm job) for about a week now
- Tried scaling to multi-node but encountered SIGBUS errors. Haven't traced down the cause.
NERSC
- Some large job failures at Cori due to not being able to find Sim_tf. Believe this was because queue was set to online instead of brokeroff after downtime..
- HammerCloud tests running on Perlmutter. Machine is down for emergency network maintenance.