SWT2_CPB:
We rebuilt four storage servers from EL7 to EL9.
There were some transfer errors caused by these rebuilds.
No data has been lost. Backups were made before rebuilds.
We checked and verified data was not lost after rebuilds were complete.
Our site experienced brief intermittent exclusion by HC.
This is a known issue to us that we believe we understand.
We have multiple servers that are changing between read-only and read-write depending on the needs of our situation with migrating and transitioning storage from EL7 to EL9.
If the number of read-only storage servers are too high, this can cause other servers that are in read/write to have a high load.
Once transitioning storage from EL7 to EL9 is complete, this issue should not occur again because less servers will be in this temporary read-only state.
GGUS-Ticket-ID: #1002282 - Jobs Mistakenly Using Home Directory
We found that jobs are using the home directory on our NAS server, causing slower performance. We have a scratch directory on worker nodes dedicated for uses such as this.
Container files are being stored and referenced here during execution of jobs.
The directories created by jobs in this area on our NAS were not cleaned up, causing a buildup of these directories.
After some discussion in the ticket, including some suggestions but also questions concerning our current working directory not being used for these container files, this area was cleaned up which did help reduce issues. The ticket was closed. However, we do not believe that this issue has been resolved. This requires having jobs stop using other directories except scratch.
OU: