gridgk03,4 were drained on 4/21 by mistake. Ivan caught and corrected this. No interruption in jobs or throughput (gridgk06,7 picked up the extra work). Things have rebalanced
Preparations are being made to migrate the Tier 1 condor nodes to use the new config we've been working on. This process should be relatively seamless. There will be a brief spike in failure rate as jobs are killed to rebuild the workers. Targeting a phased migration in batches of ~25%, with a pause after the first batch to verify jobs are flowing and completing successfully. Small scale testing so far has been good! Uptime during this whole process should remain 100% with (very) brief periods of 75% capacity
Targeting to begin next week, pending success of all the prep work (a LOT of code to verify and merge)