SWT2_CPB:
New Storage Deployment
31 files equaling 1.2 GB in size were lost during our migration on 7/23/2025 due to a rare error condition not caught by the original migration script.
We contacted Fabio Luchetti and Petr Vokac for help in reporting this.
We paused migration temporarily to test and improve our migration script in order to prevent further issues.
We have significantly improved the script, adding more safeguards and functionality, but are continuing to develop and test this in our test cluster. It is more robust and addresses the rare error condition.
We are testing this script. We will continue migration once scripts are sufficiently improved and tests are completed.
GGUS-Ticket-ID: #683657: Varnish
We gradually increased the priority of our Varnish server.
We restarted Varnish services for the new frontier, as requested by Ilija.
Communicated with Ilija concerning monitoring and other information.
We performed a one hour test with Varnish as primary proxy in CRIC on 8/5/2025 at 11:00 a.m. to 12:00 p.m. CT, then reverted Varnish back to position 1 (send priority).
We evaluated monitoring provided by Ilija and Panda job logs.
We will be performing the same test today (8/6), but for five hours instead. We will monitor and check results afterwards.
New Hardware
We received new hardware for head nodes. We plan to replace old head nodes for better performance and safety.
Declared Downtime
Declared downtime of severity warning for 8/1 and 8/2.
An external power supplier to campus performed work that temporarily removed power to the data center for one minute each day. Our UPS kept us online, and we prepared for a potential outage.
Fortunately, we did not experience an outage during this time.
UPS Issue
After the brief power losses on 8/1 and 8/2, we noticed alerts from our UPS system.
We are communicating with the UPS vendor to gather more information and recommendations going forwards to ensure we do not experience any serious issues with our UPS.
Storage Issue
One storage node is experiencing issues with one drive bay. Replacing the drive did not resolve the issue.
Currently working with the vendor for assistance in resolving this issue since it is covered by warranty.
Dark data on SCRATCHDISK
We sent the last dump of dark data 7/23. The dump file includes 30 TB of files and is created from 3 dumps. This consists of 7/8 from SWT2_CPB (all dump), 7/18 (dump Rucio data), and 7/21 (second dump from SWT2_CPB). We are waiting for dark data to be cleaned by DPM managers.
OU: