Deployment team

→ Europe/London
EVO - GridPP Deployment team meeting

EVO - GridPP Deployment team meeting

Jeremy Coles
Description
- This is the weekly DTEAM meeting - The intention is to run the meeting in EVO: http://evo.caltech.edu/evoGate/. Join the meeting in the GridPP Community area. - The phone bridge number is +44 (0)161 306 6802 (CERN number +41 22 76 71400). The phone bridge ID is 44709 with code: 4880.
Attendee ======== Jeremy Coles (Chair) Mingchao Ma (minutes) Alessandra Forti Sam Skipsey Wahid Bhimji Derek Ross James Cullen Brian Davies Daniela Bauer Dug Mcnab Richard Hellier stephen Burke Mohammad Kashif Andrew Lahiff David Colling Peter Gronbech Raja Nandakumar Experiment problems/issues ======================== - LHCb: Raja is still catching up emails. Seems to be running fine. a few disk servers down at RAL due to hardware failure; - CMS: Everything is looking ok. - ATLAS (from Graeme’s email) 1. Some HC tests seem (Friday, yesterday) seem to have validated RAL's frontier setup. In the Friday test Glasgow got 8.2Hz with FZK frontier, QMUL with 8.3Hz with RAL. 2. Oxford still down, which with this much data is not too problematic, but a 3 week outage for a T2 in the future could be a serious issue. Oxford is partially up now, passing SAM test and processing CMS jobs. Pete is still waiting for an updated report from computing staff on the problem, one chillier is down, and another one also has some problem, waiting for some work to be done this week. 3. There seems to be a problem with Cambridge's DPM not accepting US certificates (https://savannah.cern.ch/bugs/?61290, GGUS #54591). Cambridge is investigating it. 4. We will restart production at IC after their SL5 migration. SL5 migration is on the way; - Other -- camont issues with number of jobs running Camont tried to submit more jobs, but experienced some problems, only 20-30 jobs were running last week. Some sites did not accept camont jobs although usability was low. Jobs were only spreaded to RAL, Oxford and Cambridge; Glasgow did not change fair share policy recently. It seems nothing is obvious. ROC update ========= - -Update from on-duty: a quite week -- Things to follow up: Nothing to follow up -- From the EGEE ops meeting: very short meeting, nothing to follow up From the site reports (refer to agenda page for Graeme’s email): UKQCD request only concerning ECDF and Glasgow. WLCG update =========== GDB tomorrow http://indico.cern.ch/conferenceDisplay.py?confId=72046, Graeme is Tier2 representative this week. Security policies, Pilot Jobs and GGUS along with other topics will be discussed. Agenda can be found at above link. Will cover pilot framework questionnaire next week meeting when system admins are present. Glexec for SL5 is now in production. Ticket status ========== https://gus.fzk.de/download/escalationreports/roc/html/20100111_EscalationReport_ROCs.html 50491 - on hold. CMS transfers IC-RHUL. Probably jumbo frames issue. Opened in July 09*. 53364 - TCD ticket regarding fusion VO. No submitter follow up?* 53543 - Bham. Fusion file transfer. Reopened. No update from site Pete will ask site system admin. x53582 -ATLAS on hold. UCL-CENTRAL transfer problem. NAT box issue? Open Nov. Hold Dec. 53600 - Oxford Nagios. On hold pending Savannah bug fix.* No update on it from Savannah 53659 - camont submission to Bham.* Metrics ====== One area we tried to make some progress on was 4.x.14 - middleware upgrading. We'd like something here but it is tricky. An initial suggestion is that the metric per site is an AND over: WN version - CE version - SE version. Each would need to be at the WLCG recommended minimum. The Tier-2 metric is then Green if all sites meet it, Amber if 80-<100% meet it and Red if below 80%. We need to decide this week if this is an acceptable approach. Upgrading too often is not a good idea, upgrading is one of major reason of making site unstable. A better approach might be to upgrade to LCG/GridPP recommended version. A GridPP approved version mean a more stable, but some sites have to test it. How to measure the stability of a new release? For SE, a baseline recommendation version plus, GridPP storage group has its own recommendation on top of it. Probably will come back to it later. Team current focus ================ - Round the team updates (i.e. current focus) Derek: doing some planning and upgrading; reconfig lcg-ce/cream-ce; moving some server around; Will also work on SCAS and glexec; Alessandra: working to optimise ATALAS jobs and other services; will increase storage in June; start to validate values from vendors from next week; will also look into SCAS and glexec and might test it on one cluster; James: Last Friday started to see falling SAM test, now looking into it; yesterday started to look at gsexec on WNs again, will reinstall production version this week; Wahid: Looking at storm and compare Strom performance with DPM; Mingchao: will work on the security questionnaire and also organize a security training workshop at Taiwan in March, 2010 Brian: draining ATLAS disk server for the RAID5 for Tier1 ; for Tier2, limiting GridFTP transfers on DPM servers, deleting and clearing up dump files at Tier2 for ATALAs Daniela: Try to complete SL5 migration and while reading some security policy and to complete security questionnaire Dug: setup and test SCAS and glexec Richard: do some planning, specifically top level BDII service at RAL Stephen: should look at updating user guide of information systime to add more information about Glue2.0, but not clear what will happen to the user guide after EGEE III. Kashif: Looking into SCAS and glexec, has installed SCAS; will install glexec on SL5 WNs; looking into regional portal of Nagio and test Nagio package for regional portal when it releases Pete: write up quarterly report; chasing up other sites on security questionnaire and ticket; Raja: catching up things after being away Jeremy: working on GridPP4 proposal, hoping to have a first main draft version after f2f meeting on Friday; Tasks in progress ============= i) Site have been asked to complete a GridPP security questionnaire. Deadline this week. ii) Sites are going to be ticketed on publishing issues revealed by GSTAT2.0: http://gstat-prod.cern.ch/gstat/summary/GRID/GRIDPP/ iii) Staged rollout of SCAS + glexec at the T2s. Stalled at two sites. Middleware is now in production release. Oxford has setup SCAS, testing in production system, glexec is not setup yet at WNs. Glasgow has setup but will reinstall it on another server. RAL has a working SCAS server, SL4 WN working with glexect; James has SCAS server setup at the moment, a beta release glexec was setup on 3 WNs before Xmas, will setup a production glexec version on these 3 WNs in production environment. iv) Upgrading of site SEs to recommended minimum SRM releases. v) Brunel machine room move Actions ====== See http://www.gridpp.ac.uk/wiki/Deployment_Team_Action_items AOB ==== Funding for Storage workshop has been approved. Next week is site meeting. [Copy of chat window] ================== [10:57:34] Jeremy Coles joined [10:57:37] Alessandra Forti joined [10:57:37] Sam Skipsey joined [10:57:40] Wahid Bhimji joined [10:57:42] Derek Ross joined [10:57:42] James Cullen joined [10:58:45] Brian Davies joined [11:00:09] Daniela Bauer joined [11:00:12] Dug McNab joined [11:00:47] Derek Ross I think Raja's back today [11:00:59] RECORDING Mingchao joined [11:01:35] Richard Hellier joined [11:01:56] Stephen Burke joined [11:03:07] Mohammad kashif joined [11:03:07] Mohammad kashif left [11:04:32] Andrew Lahiff joined [11:05:37] David Colling joined [11:06:34] Pete Gronbech joined [11:07:50] Raja Nandakumar joined [11:20:35] Dug McNab Hi Derek, the WMS in the UK are all having problems but RAL seemed to have fixed itself. Did you guys do anything? [11:20:40] Dug McNab We are seeing Threshold for ICE Input JobDir jobs: 1500 => Detected value for ICE Input JobDir jobs /var/glite/ice/jobdir : 1514 [11:21:00] Dug McNab lots and lots of old jobs in /var/glite/ice/old [11:21:21] Derek Ross move the jobs out of the directory [11:21:28] Derek Ross seems to fix the problem [11:21:35] Dug McNab I presume this is why we are not accepting anymore jobs. ha ha, easy fix! [11:21:36] Dug McNab cheers [11:21:54] Derek Ross its a known bug fixed in the latsst/next WMS update [11:27:15] Dug McNab camont are currently limited to 100 cores at anyone time at Glasgow [11:42:42] Andrew Lahiff mic's not working [11:45:03] James Cullen I will be leaving my post a couple of weeks after the meeting, so may not be that useful to me [12:14:30] Mingchao Ma Sam, could you please update the action page and close the action, add the link to the page? [12:17:13] David Colling Sorry, I have to go ... and have been distracted here with other matters for a while. [12:17:23] David Colling left [12:23:09] Dug McNab left [12:23:09] Raja Nandakumar left [12:23:10] Mohammad kashif left [12:23:11] Wahid Bhimji bye [12:23:11] Derek Ross left [12:23:12] Andrew Lahiff left [12:23:12] Alessandra Forti left [12:23:12] Wahid Bhimji left [12:23:13] Sam Skipsey left [12:23:13] Brian Davies left
There are minutes attached to this event. Show them.
    • 11:00 → 11:20
      Experiment problems/issues 20m
      Review of weekly issues by experiment/VO - LHCb - CMS - ATLAS 1. Some HC tests seem (Friday, yesterday) seem to have validated RAL's frontier setup. In the Friday test Glasgow got 8.2Hz with FZK frontier, QMUL with 8.3Hz with RAL. 2. Oxford still down, which with this much data is not too problematic, but a 3 week outage for a T2 in the future could be a serious issue. 3. There seems to be a problem with Cambridge's DPM not accepting US certificates (https://savannah.cern.ch/bugs/?61290, GGUS #54591). 4. We will restart production at IC after their SL5 migration. - Other -- camont issues with number of jobs running
    • 11:20 → 11:30
      ROC update 10m
      ROC update *************** - Update from on-duty -- Things to follow up From the EGEE ops meeting: From the site reports: ScotGrid: 1. Glasgow and ECDF security questionnaires are in. I am not sure about Durham, though David had planned to do it last week. 2. We will enable UKQCD on our DPM this week at Glasgow and Edinburgh. WLCG update ***************** - GDB tomorrow http://indico.cern.ch/conferenceDisplay.py?confId=72046 Ticket status *************** https://gus.fzk.de/download/escalationreports/roc/html/20100111_EscalationReport_ROCs.html 50491 - on hold. CMS transfers IC-RHUL. Probably jumbo frames issue. Opened in July 09*. 53364 - TCD ticket regarding fusion VO. No submitter follow up?* 53543 - Bham. Fusion file transfer. Reopened. x53582 -ATLAS on hold. UCL-CENTRAL transfer problem. NAT box issue? Open Nov. Hold Dec. 53600 - Oxford Nagios. On hold pending Savannah bug fix.* 53659 - camont submission to Bham.*
    • 11:30 → 11:40
      Metrics 10m
      One area we tried to make some progress on was 4.x.14 - middleware upgrading. We'd like something here but it is tricky. An initial suggestion is that the metric per site is an AND over: WN version - CE version - SE version. Each would need to be at the WLCG recommended minimum. The Tier-2 metric is then Green if all sites meet it, Amber if 80-<100% meet it and Red if below 80%. We need to decide this week if this is an acceptable approach.
    • 11:40 → 11:48
      Team current focus 8m
      - Round the team updates (i.e. current focus)
    • 11:48 → 11:58
      Tasks in progress 10m
      i) Site have been asked to complete a GridPP security questionnaire. Deadline this week. ii) Sites are going to be ticketed on publishing issues revealed by GSTAT2.0: http://gstat-prod.cern.ch/gstat/summary/GRID/GRIDPP/ iii) Staged rollout of SCAS + glexec at the T2s. Stalled at two sites. Middleware is now in production release. iv) Upgrading of site SEs to recommended minimum SRM releases. v) Brunel machine room move
    • 11:58 → 12:03
      Actions 5m
      See http://www.gridpp.ac.uk/wiki/Deployment_Team_Action_items
    • 12:03 → 12:04
      AOB 1m