SWT2_CPB:
EL9 Migration
We are currently focusing on deploying new storage, migrating old data, and retiring old storage, but we do have new modules to test once we are done using the test cluster for new storage deployment. Later, we will work toward implementing and migrating additional server types to EL9.
New Storage Deployment
We started migrating data from one MD3460 storage to one of the new R760xd2 storage servers. We are carefully monitoring this, checking data integrity, and improving the process as we go.
All twelve new R760xd2 storage servers are installed with EL9. We are focusing on making sure we safely migrate the first server for understanding before using the other storage servers.
GGUS-Ticket-ID: #681994: Enable network monitoring
We met the requirements and received approval from campus network security for enabling network monitoring on a web server.
We converted the SNMPv2c script provided by Shawn to use SNMPv3 instead. I shared this with Shawn in case other sites ever need to use it.
Our network monitoring is now enabled.
GGUS-Ticket-ID: #681997: Enable BGP Tagging
We worked with campus networking to get approval, they implemented it, then we checked and confirmed with Edoardo to verify it is correct.
GGUS-Ticket-ID: #683657: Implement Varnish
We ran a test in the test cluster according to Ilija's documentation.
We created a new Puppet module for Varnish, tested it, then converted gk02 to a Varnish server. It is fully configured, running, and appearing in Ilija’s Varnish monitoring dashboard.
We want to run jobs from the test cluster using gk02 separately from production PQ if possible before placing gk02 frontier proxy higher on the priority list in CRIC.
GGUS-Ticket-ID: #1000094: Reallocating Scratchdisk
We have adjusted the max size of SCRATCHDISK from 500 TB to 300 TB.
Certificate Issue
Result HTTP 401: Authentication Error after 1 attempts” 1.64 K for 24 h for SWT2 - dst DDM, and all sites src DDM
We noticed low transfer efficiency to RU, FR, and IT with these errors:
We discovered an issue with how we update our certificates. We updated the osg-ca-certs package on our static repo to version 1.136 and are considering changes to avoid doing this manually.
We are currently using OSG 23, but do not want to update this as it will change the version of packages that may cause us issues. We updated our current static repo’s osg-ca-certs package only for now.
After performing this change, we did not see a full recovery in transfer efficiency, but we may need to monitor and give it more time to reflect.
Hardware
We purchased new hardware to upgrade head nodes. Waiting for shipment.
We extended the warranty on several R740xd2s that were going to expire mid-July of this year.
OU: