SWT2_CPB:
In terms of new storage and EL9 migration, continuing to work with DDM ops to clear up remaining issues that need to be configured for the test cluster in order to perform additional tests for both EL9 migration and new storage.
Set up storage dumps and provided requested info to DDM ops for test cluster setup in CRIC.
We have been performing manual tests using commands and are now testing/improving migration scripts.
We are continuing to adjust the storage module for improvements. Adjusting and testing partition configurations.
Have additional modules created and are waiting to test with the test cluster. In the meantime, we review these and are making improvements where we can.
Improved our internal monitoring to include both CEs.
Created and are developing new internal alerting using Zabbix. Currently developing in the test cluster.
Thanks to Ivan’s help, we switched from 16 cores back to 8 core jobs. Majority of our running slots have been filled with 8 core jobs recently.
Set CE to remove completed jobs after ten days to prevent stuck jobs from causing issues as suggested by experts.
Created alerts for jobs that are completed and have been on the CE for more than six days.
Adjust and testing parameters for both of our CE with the consultation of the harvester team to better improve load balancing of CEs.
Added additional nodes to the test cluster in order to test Varnish and other Puppet modules.
Gathering information and communicating with Dell to purchase new hardware to replace head nodes.
GGUS tickets - Concerning BGP tagging and network monitoring, tried reaching out to campus networking for results of their discussions, but have not heard a response. Plan on contacting other members to schedule a meeting.
Request to deploy IPv6 on CEs and WNs at WLCG sites - We have configured an IPv6 address on gk01 and test CE. HTCondor-CE service has to be restarted for it to take effect. Because of this, we will wait until gk10 is completely drained during our load balancing tests, configure IPv6, then restart the service. We do not see a reason to add IPv6 to WNs at this time, because they are isolated to an internal network.