- ROD team update
- EGI operations update (thanks to Stuart P)
"EMI-1 software release:
WMS: On the 17th there was an emergency release of WMS 3.3.1 - addresses a vulnerability
CREAM: Release expected on 23rd to fix infinite loop with argus integration.
Storm: Not now expected to be released for the 23rd June; delayed. (Might make next update release on 7th July ?)
Next major update is expected on the 7th July, for components that get released to certification before 27th June. If they don't make it then it will be 2 weeks later. Detailed list of planned release in: https://www.egi.eu/indico/getFile.py/access?contribId=0&resId=0&materialId=slides&confId=495. Includes L&B, WMS and Argus updates.
Some discussion on tickets, and tracking of these - in particular who is responsible for making sure that tickets are also cc'd to the Release managers. Mario will be updating procedures to make this clear (it's the staged rollout managers).
UMD/Staged rollout: Useful table at https://wiki.egi.eu/wiki/Agenda-20-06-2011. Moving UMD release announcement from 4th to 11th July
Operational Documentation: On how to migrate a service preserving it's state. A survey was done of what is present, and what more is needed; work is underway. Consideration of this issue for ARC and Unicore. This will be added as a requirement for EMI."
- Nagios status
Tier1 comments from this last week.
============================
Some problems with the CMS Castor instance last week (Wednesday). Load issues in the LSF scheduler that is used internally within Castor. 90minutes unscheduled downtime declared.
Thursday: The DB team reported that intervention to update an Oracle and a system parameter on remaining Castor database nodes went OK. Although too early for confirmation this should fix a problem to do with gathering the database statistics.
Over the weekend we had a problem with LHCb SRMs.
We also had a problem with CERN information missing from our top BDIIs.
There have been a couple of disk server problems in the last week. (One for Alice, one for LHCb). No data loss
We are looking to apply a partial update to Castor in order to prepare for the higher capacity "T10KC" tapes. This is likely to take place on Tuesday 5th July during the LHC technical stop. We are also rolling out a newer version of the Castor clients across the worker nodes as soon as practical.
At LHCb's request we have increased the maximum wall clock limit on our 6GB queue by 20% (to 120 hours).
- Security update
Torque vulnerability follow-up
-- T2 issues
--- Accounting (http://www3.egee.cesga.es/gridsite/accounting/CESGA/egee_view.php). Durham?
--- http://espace.cern.ch/WLCG-document-repository/ReliabilityAvailability/Tier-2/2011/Tier2_reliab_05-2011.pdf
- Check on UCL (41%:28%); EFDA-JET (73%:49%) and BHAM (87%:87%)
- HEPSYSMAN
-- 30th June - 1st July (RAL): http://hepwww.rl.ac.uk/sysman/June2011/agenda.html