Operations team & Sites

Europe/London
Description
- This is the biweekly Ops-team & sites meeting - The intention is to run the meeting in EVO: http://evo.caltech.edu/evoGate/. Join the meeting in the GridPP Community area. - The phone bridge number is +44 (0)161 306 6802 (CERN number +41 22 76 71400). The phone bridge ID is 77907 with code: 4880. Apologies: Mark M; Tier-1 folk (away day), Wahid
    • 11:00 11:20
      Meetings & updates 20m
      - ROD team update - EGI operations update (thanks to Stuart P) "EMI-1 software release: WMS: On the 17th there was an emergency release of WMS 3.3.1 - addresses a vulnerability CREAM: Release expected on 23rd to fix infinite loop with argus integration. Storm: Not now expected to be released for the 23rd June; delayed. (Might make next update release on 7th July ?) Next major update is expected on the 7th July, for components that get released to certification before 27th June. If they don't make it then it will be 2 weeks later. Detailed list of planned release in: https://www.egi.eu/indico/getFile.py/access?contribId=0&resId=0&materialId=slides&confId=495. Includes L&B, WMS and Argus updates. Some discussion on tickets, and tracking of these - in particular who is responsible for making sure that tickets are also cc'd to the Release managers. Mario will be updating procedures to make this clear (it's the staged rollout managers). UMD/Staged rollout: Useful table at https://wiki.egi.eu/wiki/Agenda-20-06-2011. Moving UMD release announcement from 4th to 11th July Operational Documentation: On how to migrate a service preserving it's state. A survey was done of what is present, and what more is needed; work is underway. Consideration of this issue for ARC and Unicore. This will be added as a requirement for EMI." - Nagios status Tier1 comments from this last week. ============================ Some problems with the CMS Castor instance last week (Wednesday). Load issues in the LSF scheduler that is used internally within Castor. 90minutes unscheduled downtime declared. Thursday: The DB team reported that intervention to update an Oracle and a system parameter on remaining Castor database nodes went OK. Although too early for confirmation this should fix a problem to do with gathering the database statistics. Over the weekend we had a problem with LHCb SRMs. We also had a problem with CERN information missing from our top BDIIs. There have been a couple of disk server problems in the last week. (One for Alice, one for LHCb). No data loss We are looking to apply a partial update to Castor in order to prepare for the higher capacity "T10KC" tapes. This is likely to take place on Tuesday 5th July during the LHC technical stop. We are also rolling out a newer version of the Castor clients across the worker nodes as soon as practical. At LHCb's request we have increased the maximum wall clock limit on our 6GB queue by 20% (to 120 hours). - Security update Torque vulnerability follow-up -- T2 issues --- Accounting (http://www3.egee.cesga.es/gridsite/accounting/CESGA/egee_view.php). Durham? --- http://espace.cern.ch/WLCG-document-repository/ReliabilityAvailability/Tier-2/2011/Tier2_reliab_05-2011.pdf - Check on UCL (41%:28%); EFDA-JET (73%:49%) and BHAM (87%:87%) - HEPSYSMAN -- 30th June - 1st July (RAL): http://hepwww.rl.ac.uk/sysman/June2011/agenda.html
    • 11:20 11:40
      Experiment problems/issues 20m
      Review of weekly issues by experiment/VO - LHCb - CMS - ATLAS - Other - Experiment blacklisted sites - Experiment known events affecting job slot requirements -- The European Physical Society’s High Energy Physics conference will be held in Grenoble from 21 to 27 July; the Lepton-Photon conference at the Tata Institute in Mumbai from 22 to 27 August.
    • 11:40 11:55
      Site roundtable (glexec focus) 15m
      - Note the the gLExec deployment page has been updated: https://twiki.cern.ch/twiki/bin/view/LCG/GlexecDeployment. (Glasgow's SCAS/glexec wiki entry http://www.scotgrid.ac.uk/wiki/index.php/Glasgow_GLite_gLExec_installation_and_configuration). Checking the status: https://samnag023.cern.ch/nagios/cgi-bin/status.cgi?hostgroup=United+Kingdom&style=detail
    • 11:55 12:00
      Core ops team task areas 5m
      - Ahead of the team discussion of the task areas it would be useful to work with individuals in the core team to gather additional ideas and to develop the descriptions in the wiki. Areas (http://www.gridpp.ac.uk/wiki/Category:GridPP_Operations) are: - Staged rollout - On-duty coordination - Ticket follow-up - Regional tools (Kashif+?) - Documentation - Security (Alessandra/Jeremy) - Monitoring - Accounting - Core grid-services - Wider VO-services There will be other areas like incident follow-up, coordination with other grids (Stuart?) that may appear explicitly.
    • 12:00 12:01
      AOB 1m
      - WLCG workshop http://indico.desy.de/conferenceDisplay.py?confId=4019. Please submit travel requests and make bookings early. - Plug for: " tomorrow’s [Tier-1 liaison] meeting to be held at 13:30 via EVO only. * Meeting URL: http://evo.caltech.edu/evoNext/koala.jnlp?meeting=eleBevvaveanaIaeIl * Phone bridge ID: 8 3566 Experiments Liaison Meeting: http://www.gridpp.ac.uk/wiki/RAL_Tier1_Experiments_Liaison_Meeting" - And.... Lustre Workshop will be held at QMUL on 14 July (note new date). Please see http://www.lustreusergroup.org/ to sign up and for details of the programme. Still a little space if someone wants to present something.