Deployment team & sites

→ Europe/London
EVO - GridPP Deployment team meeting

EVO - GridPP Deployment team meeting

Jeremy Coles
Description
- This is the weekly DTEAM meeting - The intention is to run the meeting in EVO: http://evo.caltech.edu/evoGate/. Join the meeting in the GridPP Community area. - The phone bridge number is +44 (0)161 306 6802 (CERN number +41 22 76 71400). The phone bridge ID is 44709 with code: 4880.
Minutes
    • 11:00 → 11:20
      Experiment problems/issues 20m
      Review of weekly issues by experiment/VO - LHCb - CMS - ATLAS T1/T2/T3 jamboree next Tuesday: http://indico.cern.ch/conferenceDisplay.py?confId=76900 - Other - More NFS issues? - Blacklisted sites? Other VOs: - e-NMR - superbvo http://www.gridpp.ac.uk/wiki/GridPP_approved_VOs
    • 11:20 → 11:30
      ROC update 10m
      ROC update *************** - Update from on-duty -- Things to follow up Note that the GridPP VOMS server cert changes next week: the server certificate on voms.gridpp.ac.uk is due to expire on the 11th of February. The new certificate will be installed on the 1st of February between 8 and 8.30 am UTC. Links to the new certificate (pem / rpm / yum) can be found on http://www.gridpp.ac.uk/wiki/Instruction_for_VO_administrators#Current_certificates Tier-1 status ********** - Some new problems for the Tier-1.....http://www.gridpp.rl.ac.uk/blog/2010/02/01/update-on-tier1-castor-problem-and-3d-migration/ WLCG update ***************** - Machine start-up still scheduled for mid-February. Current talk of an extended run ahead of a long shutdown for upgrades needed for higher energy collisions. Ticket status *************** https://gus.fzk.de/download/escalationreports/roc/html/20100201_EscalationReport_ROCs.html 50491 - on hold. CMS transfers IC-RHUL. Probably jumbo frames issue. Opened in July 09****. 53349 - on hold. Bristol. Publishing vast amount of storage. Opened in November**. 53363 - Lancaster +2 (split). Fusion issue with .lsc file? Last update 08/01. 53364 - TCD ticket regarding fusion VO. No submitter follow up?** Wait on submitter. Close?** 53598 - ATLAS T1. On hold. Channel load change request. wait for data to test? 53600 - Oxford Nagios. On hold pending Savannah bug fix.****
    • 11:30 → 11:35
      Accounting/HEPSPEC06 review 5m
      - Situation with APEL - Checking the published data: http://www3.egee.cesga.es/gridsite/accounting/CESGA/egee_view.php. Currently stops just after mid-January. Database problems -> the impacts are: - Accounting portal not updated - SAM tests not updated Sites should still be able to publish normally and there is no need for them to do anything. Downtimes on APEL have been declared, and operators should be warned not to open GGUS tickets against APEL failures until this is resolved. - Are any sites now running NOT using the HEPSPEC06 measured data to give the KSI2K rating?
    • 11:35 → 11:42
      SCAS & glexec questionnaire responses 7m
      - Where next... The questionnaire (also put on last week's agenda for those who think this summary is familiar) Does your site policy allow the use of MUPJ by the LHC experiments you support? (no/depends/yes) • Yes 16/17 • Depends 1/17 • No • Does your site policy support the use of glexec in setuid mode? (no/allow/require) • No • Allow 14/17 • Require 3/17 • Does your site policy support the use of glexec in log-only mode? (no/allow/require) • No 4/16 • Allow 12/16 • Require To smooth the deployment and use of the critical components involved, an exceptional treatment of the following case may be desirable until sufficient stability has been demonstrated across WLCG: • When glexec returns an internal error (e.g. SCAS/Argus/GUMS temporarily unavailable), does your site policy allow the pilot to continue and run the payload itself? (no/depends/yes) • No 4/16 • Depends 10/16 • Yes 2/16 - Next steps are for the technical forum to review WLCG feedback overall. - To come back to you for clarification where required - Put forward conclusions/recommendations ahead of March GDB In the background testing of SCAS/glexec continues - We have a few instances in production - Testing has so far produced positive feedback - Further supporting sites requested CREAM: - Seems to be working well at several sites (feedback?) - The EGEE target was for all sites to deploy CREAM in parallel to LCG CE by end of March - We would like some more sites to deploy this month.
    • 11:42 → 11:56
      Current site problems 14m
      - Site-by-site review of status/concerns/problems
    • 11:56 → 12:03
      Security 7m
      Chance to discuss the questionnaire, current (ongoing) incidents, best practice....
    • 12:03 → 12:04
      AOB 1m