Last week our "OtherVO" gatekeeper suddently stopped working. I noticed this first on April 19 when our space-usage.json file stopped updating dating to mid-afternoon on April 18. Then I noticed that the CE jobs coming in were going into Hold, and never emerging to become real batch system jobs. Basically, all services were timing out.
Chased this hard, even going so far as to totally rebuild that gatekeeper (which is good in the sense that gram should now be gone totally from it). Nothing worked until I moved it temporarily from the grid.umich.edu domain to the aglt2.org domain, and saw both gfal-copy and condor_ce_ping begin to work.
Subsequently found that the UM subscribes to an IPS Service (Intrusion Prevention System) that had distributed a complete blockage on TSL/SSL traffic. Our subnet was scanned to their satisfaction, it was white-listed, and the gate-keeper then came back online.
Beware of the possibility of a similar situation at your home Universities. It has also affected the MWT2 and Ligo collborators at the UM.
A dCache OOM issue hit us when the number of multiple srm transfers reached levels above what we've ever seen here (a peak of 18k/10mins was observed a few days ago in srmwatch) in conjunction with just more memory being used anyway with dCache 3.0.11. We increased the dCacheDomain java instance memory from 2/4gb to 4/8gb and the issue resolved.
Beyond this services are stable. Today we will complete the transition to the last of our new N2048 switches for public 1Gb NIC attachment.
For queue ANALY_AGLT2_TEST_SL6-condor, the wansinklimit and deprecate_old_mover flags are now properly set (as of 1:30pm today).