Operations team & Sites

Europe/London
EVO - GridPP Operations team meeting

EVO - GridPP Operations team meeting

Description
- This is the biweekly ops & sites meeting - The intention is to run the meeting in EVO: http://evo.caltech.edu/evoGate/. Join the meeting in the GridPP Community area. - The phone bridge number is +44 (0)161 306 6802 (CERN number +41 22 76 71400). The phone bridge ID is 126540 with code: 4880. Apologies:
Minutes
    • 11:00 11:20
      Experiment problems/issues 20m
      Review of weekly issues by experiment/VO - LHCb - CMS - ATLAS ATLAS report for week 14 – 21 February 2012 ====================================== - ATLAS LFC has successfully migrated from RAL to CERN. Cloud was set offline 2012-02-13 20:48:38 and set back online on 2012-02-16 11:25:15 - All UK sites have availability >90% for analysis in January: http://atladcops.cern.ch:8000/drmon/crmon_TiersInfo.html - MC11 campaign is practically finished, MC12 campaign should start in the beginning of March (8TeV) - UCL is in DT (DPM upgrade), blacklisted and prod queues are offline, https://ggus.eu/ws/ticket_info.php?ticket=79276 - closed - Brunel: changing DPM head node. Storage has been replaced both in ToA and Schedconfig after all the datasets were deleted over the weekend. Test jobs were submitted. Transfers started to fail because of permission problem. Blacklisted. https://ggus.eu/ws/ticket_info.php?ticket=79387 __________________________________________________ - Other - T2K
    • 11:20 11:40
      Meetings & updates 20m
      - ROD team update - EGI ops - Nagios status - Tier-1 update - Security update - T2 issues - General notes There was a GDB the week before last: https://indico.cern.ch/conferenceDisplay.py?confId=155065. Areas covered: For those interested in the Technical Evolution Group summaries presented on the Tuesday look here: https://indico.cern.ch/conferenceDisplay.py?confId=158775. Site dashboards: http://dashb-siteview.cern.ch/templates/siteview/index.html EMI status: Planned updates http://bit.ly/all_active_tasks. See also discussion on end of security patching. LHCOPN & LHCONE: http://lhcone.net. IPV6 CERN-KIT. Perfsonar dashboard improved https://perfsonar.usatlas.bnl.gov:8443/exda/?page=25&cloudName=LHCOPN. And the LHCONE dashboard (15 members) https://130.199.185.78:8443/exda/?page=25&cloudN ame=LHCONE. Security - for interest. Storage accounting: Definition of usage record for storage. Going OGF UR 1 route. SE sensors in place by May. Publishing around June. RFC/SHA-2 proxies: IGTF would like CAs to move from SHA-1 to SHA-2 signatures ASAP, to anticipate concerns about the long-term safety of the former. For WLCG this implies using RFC proxies instead of the Globus legacy proxies in use today. 10 months to get ready for RFC and SHA-2. There are various pieces of middleware and experiment-ware that need to be made ready for SHA-2 or RFC proxy support. For EMI products the current time line is the EMI-2 release in April/May. Pete's report will appear here: https://www.gridpp.ac.uk/wiki/GDB_reports. Representation for the March meeting!? The focus is expt. operations updates: https://indico.cern.ch/conferenceDisplay.py?confId=155066. - Tickets The NGI: https://ggus.eu/ws/ticket_info.php?ticket=78991 The ticket documenting the removal of the email address field from the gridpp voms server's certificate. This ticket has evolved to include the fate of the e-mail address field in the host certificates. RAL Tier-1 https://ggus.eu/ws/ticket_info.php?ticket=79283 LHCB were having job problems on the CREAM CEs, requiring a purging of lhcb jobs. This was done and it looks like problems don't persist. Can this ticket be closed? https://ggus.eu/ws/ticket_info.php?ticket=77026 RAL BDII problem noticed by Chris. I need to check up on this, repeating Chris' original simple test would be a good start. UCL: https://ggus.eu/ws/ticket_info.php?ticket=79276 Atlas blacklisted the site due to a SRM downtime, which itself caused some confusion (the downtime was originally set as "warning" rather then "outage", and then after the mistake was rectified the gocdb defaulted the downtime back to "warning" when Ben extended the length of the downtime - that's a feature to look out for). MANCHESTER: https://ggus.eu/ws/ticket_info.php?ticket=78776 Ops tests failing due to "lack of space". Kashif mentions ticketing the dpm developers or somehow figuring out a workaround. Sounds like we need to refer it to the storage group. QMUL: https://ggus.eu/ws/ticket_info.php?ticket=77959 Atlas deletion errors on the QMUL Storm SE. Has Chris had a chance to come back to it after his holiday?
    • 11:40 12:05
      Site roundtable 25m
      - What's going on across the sites
    • 12:05 12:10
      AOB 5m
      - Request from Mark M: site traffic figures for December, January and part of this month please.