SWT2_CPB:
We implemented a second XRootD redirector for redundancy and rebuilt our EL7 XRootD redirector to EL9.
We performed this in stages: Tested this in the test cluster, implemented the second EL9 redirector, removed the EL7 redirector, transitioned the EL7 redirector to EL9, then added it back into production.
We had a configuration error that led to errors and HC blacklist on 7/10. We realized that we needed to restart the services on storage servers and that the configuration on one of the XRootD redirectors was incorrect. We restarted the XRootD service on these nodes and corrected the configuration.
Nodes in SWT2_CPB_K8S nodes are using EL7 admin node for DNS lookups. Because the newer version of XRootD uses hostname instead of IP address for storage and EL7 admin did not have EL9 storage in DNS, jobs failed on these worker nodes on 7/15. We changed DNS on admin to fix this issue.
EL7 admin node updated DNS changes on 7/20, removing the fix we performed on 7/15, which started causing the same errors as before. We fixed DNS again and added a more permanent solution. We are discussing the SWT2_CPB_K8S queue.
Once these issues have been fixed, we have not seen issues.
GGUS-Ticket-ID: #1003053 - IGTF CRLs & fetch-crl SHA1 signature validation (SWT2_CPB)
We removed gk05 from the DNS round robin.
We rebuilt gk05 and are using it for testing of other tasks, such as testing a potential network upgrade.
We plan to inspect this server more thoroughly and will add it back to the DNS round robin to test for the same errors.
OU: