Deployment team

→ Europe/London
EVO - GridPP Deployment team meeting

EVO - GridPP Deployment team meeting

Jeremy Coles
Description
- This is the weekly DTEAM meeting - The intention is to run the meeting in EVO: http://evo.caltech.edu/evoGate/. Join the meeting in the GridPP Community area. - The phone bridge number is +44 (0)161 306 6802 (CERN number +41 22 76 71400). The phone bridge ID is 44709 with code: 4880.
Attending: Alessandra, James, Jeremy, Sam, Dug, Raja, Wahid, Mingchao, Graeme Experiment problems/issues ==== - LHCB Nothing running on the Tier 2's at present. Awaiting restart in Feb. NFS experiment software issues still occur. Not necessarily at Tier 2's now. - ATLAS Generally Running fairly smooth. Manchester disk issues - back online. Glasgow disk issues - RAID card failed. Important MiniBias data recovered. Now being re-subscribed from local off-grid copy back to Glasgow. ECDF - disk server failure. Production data lost. QMUL Software Issues. Mis-configuration in pilot factory caused issues. Allessandro De Salvo started fresh installs but ATLAS UK did not know to stop sending pilots. Confusion over CE's now fixed and things installing correctly. GLEXEC pilot branch needs to be merged into trunk and testing can begin at sites with GELXEC installed. - CMS no representation. - Experiment blacklisted sites not discussed. ROC update ========== - RCOD tools Pete and Kashif to report on this from Lyon. - EGEE Ops Meeting Nothing to report. Very short meeting. - WLCG update GDB Selection - 10th Feb Alessandra Tickets ======= 53834 - ECDF - Wahid reported that they will respond to the ticket and set a date for upgrade. December WLCG - ops performance & discounting SAM results ========== UK Tier 2 Results looking good. Detailed site results on page 7 and 8. LDN - UCL Central/ RHUL possible issues. SouthGrid - Oxford - air con issue, Birmingham Alessandra and Sam described DPM information publishing GGUS ticket: 54818. The dpm-listspaces injunction to calculate free space for publishing (in --gip mode) appears to double count the "unavailable" free space due to RDONLY filesystems, resulting in the published "free" space for the system being reported as 0. This caused site failures for a period of time that was not a site issue but a middleware bug. Acton: Jeremy to find out from Alberto how the report is derived and how it can be corrected. Jeremy : Do we discount the SAM tests for this period? Jeremy: Do we adjust the raw data? Sam: Do we annotate the raw data? We know the SE and CE tests that would have failed. Can these be amended? Quarterly reports ========== https://www.gridpp.ac.uk/deployment/status/reports/reports.html - Alessandra - NorthGrid : some discussion on Lancaster. - Graeme - ScotGrid: News - Good Engagement with Atlas at Glasgow, ECDF now running better allowing opportunistic use. Biomed support removed on WMS at Glasgow as there was no response to GGUS tickets about over-quota users. MUPJ questionnaire results to note ========== The results now to be fed by to Martin Litmaath. He will summarise for the technical forum and then conclusions and recommendations will be presented to GDB. UK T2 representative for OPN discussions ========== Brian has been suggested to be put forward for a UK T2 representative. Someone with experience with an OPN. CHAT ==== [11:00:24] Jeremy Coles joined [11:00:26] Sam Skipsey joined [11:00:28] Wahid Bhimji joined [11:01:02] Alessandra Forti joined [11:01:30] Dug McNab I currently don't have my head set, so my apologies in advance if I speak and sound terrible . [11:02:38] Alessandra Forti I can hear but not speak [11:03:56] Raja Nandakumar joined [11:04:17] Wahid Bhimji you probably got the easy minute taking day Dug ! [11:04:44] Mingchao Ma joined [11:06:09] Dug McNab i know [11:46:04] Graeme Stewart joined [11:46:17] Dug McNab we received it as an EGEE broadcast [11:46:59] Dug McNab quick question, do we need to update the certificate or will the LSC filestake care of it? ACTIONS ======= Jeremy's actions still outstanding gstat2.0 checking still outstanding LHCB VOMS server action closed. AOB ======= Manchester have moved their two VOMS servers to new, more powerful hardware. Savannah ticket opened for memory leaks. Manchester VOMS server certificate requires upgrading at sites supporting VO's hosted on Manchester's VOMS.
There are minutes attached to this event. Show them.
    • 1
      Experiment problems/issues
      Review of weekly issues by experiment/VO - LHCb - CMS - ATLAS "QMUL had to move their software area to a different disk and that apprently requires a complete reinstall of the ATLAS software as it seems to hardcode the mount point in various places. This was requested on the 18th (Ticket 54720). Now instead of doing so, different bits of Atlas keep issueing tickets to QMUL e.g. 54795 and 54711 which they then seem to ignore. I was hoping to get somebody in Atlas to take charge of this issue, to get it fixed." - Other - Experiment blacklisted sites - Site performance
    • 2
      ROC update
      ROC update *************** - Update from on-duty -- Things to follow up -- Input for the RCOD tools meeting later this week From the EGEE ops meeting: From the site reports: WLCG update ***************** - The next GDB is on 10th Feb (pre-GDB is an ATLAS jamboree day) - Then 24th March - 12th May - 9th June - 14th July* - 9th August* Ticket status *************** https://gus.fzk.de/download/escalationreports/roc/html/20100125_EscalationReport_ROCs.html 50491 - on hold. CMS transfers IC-RHUL. Probably jumbo frames issue. Opened in July 09***. 53349 - on hold. Bristol. Publishing vast amount of storage. Opened in November*. 53363 - 53364 - TCD ticket regarding fusion VO. No submitter follow up?** Wait on submitter. Close?* 53600 - Oxford Nagios. On hold pending Savannah bug fix.*** 53659 - camont submission to Bham.** Waiting on site. 53834 - ECDF LCG-CE out-of-date. Opened early December. Waiting on site.*
    • 3
      December WLCG - ops performance & discounting SAM results
      "Find below the draft of the Reliability and Availability Report for the WLCG Tier-2 sites for last month. https://twiki.cern.ch/twiki/bin/viewfile/LCG/SamMbReports?filename=Tier2_Reliab_200912.pdf Please verify your data and send your comments and corrections to lcg.office@cern.ch before Friday 29 January 2010. This report will then be published on the WLCG web site and reported to the WLCG Overview Board". - A question arose last week about site availability where the result was caused by a middleware problem. Previously for site testing it was up to the deployment team to decide if a given accounting period results could be adjusted so this question will end up with us again. What criteria would we apply prior to make a formal request for a change on the SAM repository?
    • 4
      Quarterly reports
      To go through any issues... https://www.gridpp.ac.uk/deployment/status/reports/reports.html
    • 5
      MUPJ questionnaire results to note
      • Does your site policy allow the use of MUPJ by the LHC experiments you support? (no/depends/yes) • Yes 16/17 • Depends 1/17 • No • Does your site policy support the use of glexec in setuid mode? (no/allow/require) • No • Allow 14/17 • Require 3/17 • Does your site policy support the use of glexec in log-only mode? (no/allow/require) • No 4/16 • Allow 12/16 • Require To smooth the deployment and use of the critical components involved, an exceptional treatment of the following case may be desirable until sufficient stability has been demonstrated across WLCG: • When glexec returns an internal error (e.g. SCAS/Argus/GUMS temporarily unavailable), does your site policy allow the pilot to continue and run the payload itself? (no/depends/yes) • No 4/16 • Depends 10/16 • Yes 2/16 So MUPJ is widely supported. Glexec can on the whole be used in setuid mode and over half the sites can run in log-only mode. Running when glexec returns an error in most cases depended on whether the experiments were accurately keeping logs of the jobs submitted. There were many additional comments which the technical forum will likely have to iterate on with the site admins directly.
    • 6
      UK T2 representative for OPN discussions
      - The LHCOPN group are expanding discussions to include T2s - They want a representative from the UK who has some level of technical understanding in this area - Pete and Robin to be kept in the loop
    • 7
      Actions
      See http://www.gridpp.ac.uk/wiki/Deployment_Team_Action_items
    • 8
      AOB