The 'fetch-crl' command went missing (https://stfc.atlassian.net/l/cp/01CAs1Gv) and many CRLs expired on Sunday about 5pm. Failures in the SAM tests on all CEs, Echo webdav (not xrootd), Antares xrootd (not webdav) and all the AAA machines. Echo webdav was partially fixed on Sunday around midnight, with intermittent failures after that. Everything was fixed before lunch on Monday.
About 750TB of CMS data for Antares was in backlog due to a CMS WM bug. Rules that should have been distributed over the previous 1 month were created for approval on Wednesday afternoon and about 2000 of them were approved on Thursday, with the data starting to hit Antares from Thursday night. On Friday afternoon the remaining 2500 rules were approved. Some issues, for the record:
As a result of the above, SAM status was red on Friday, Sunday, Monday, Tuesday. Katy removed CMS from drain on Monday after the CRL was fixed. Interestingly, because other VOs started draining before CMS, CMS picked up many slots (35k at max!) before draining from that high value. Job performance remained good during this time. CMS went into drain again this morning (Wed), Katy removed the drain status before lunch.
The 'production' Shovler instance has been switched from Cloud to VMWare. VMWare is more resilient for a production service.
IPv6 connectivity of the AAA machines. A ticket has been sent to DI, which Katy cannot view but requested the service desk to progress it.