SWT2_CPB:
We rebuilt four R740 storage servers from EL7 to EL9 while preserving data.
After the rebuilds, we verified the data and it appears to have been preserved. We have a temporary backup of the data in case of any data loss.
The servers have been returned to production, and we have not observed any issues so far.
We are currently creating backups of additional R740 storage servers for migrating them from EL7 to EL9.
So far, we have not observed any issues with transfers related to these rebuilds.
We are still experiencing unavailable logs for most failed jobs with SIGTERM error.
We changed the default walltime limit from 48 to 78 hours on both of our Condor-CE.
We changed the KillWait value in our Slurm configuration from 5 to 10 minutes.
A new second backup Varnish server, gk12.atlas-swt2.org, was added to our site’s proxy list in CRIC.
Its priority was set to the second position in case our main Varnish server fails.
We experienced roughly 156 stage-out errors and 133 stage-in errors on 3/17.
We noticed three storage servers had very high load.
We are still investigating this issue.
The number of errors seems to be mostly associated with these storage servers.
OU: