UKI Monthly Operations Meeting (TB-SUPPORT)

Europe/London
EVO - GridPP Deployment team meeting

EVO - GridPP Deployment team meeting

Jeremy Coles
Description
- This is the monthly UKI meeting - The intention is to run the meeting in EVO: http://evo.caltech.edu/evoGate/. Join the meeting in the GridPP Community area. - The phone bridge number is +41 22 76 71400. The phone bridge ID is 482161 with code: 4880.
Minutes
    • 11:00 11:15
      Site readiness & availability 15m
      SAM tests: http://pprc.qmul.ac.uk/~lloyd/gridpp/samtest.html - recent CE warnings at number of sites? - QMUL CE issues (SGE & hardware) - Lancaster (recent CE and SE/SRM test failures) - Birmingham 84% last 24hrs? UK tests: http://pprc.qmul.ac.uk/~lloyd/gridpp/uktest.html - RAL-T1 running most jobs but faiing them at 95% level! - QMUL & RHUL have 25% falure rate - Bristol 12% failure rate ATLAS tests: http://pprc.qmul.ac.uk/~lloyd/gridpp/atest.html - Recent problems at TRINITY; IC (HEP; LeSC and HPC); some of the RHUL and QMUL clusters; ECDF; Oxford and RAL-T1 LHCb tests; http://pprc.qmul.ac.uk/~lloyd/gridpp/lhcb_samtest.html - Sites with problems highlighted by tests above except Durham. Accounting: http://www3.egee.cesga.es/gridsite/accounting/CESGA/egee_view.php - Lancaster (end of July) - Birmingham (start of August) GridMap: http://gridmap.cern.ch/gm/ - Concern about the overall picture leading up to 10th September. - Currently "down": UCL-CENTRAL; ECDF; RALPP; OX-HEP; IC-HEP - Currently "degraded": RAL-T1; QMUL; Brunel & RHUL. - But still UKI is running many jobs....
    • 11:15 11:25
      Experiment problems/issues 10m
      Review of issues by experiment/VO Latest WLCG ops discussions are here: https://twiki.cern.ch/twiki/bin/view/LCG/WLCGDailyMeetingsWeek080825. - LHCb Issues : RAL down today (till lunchtime) Because of an error in one of the LHCb install projects, the tests are now reduced to reflecting the results of the OS of the sites. Roma and Stuart need to fix the SAM tests. - CMS CRUZET-4 has started - ATLAS Most pressing news for us is that CASTOR at RAL has been down for 10 days now. No data is coming to RAL and nothing which is there is accessible. Production is halted in the UK and many T2s lost a lot of perfectly good job outputs. RAL said they would come back today ... worrying that we are down 2 weeks before data taking. Dowtime just extended due to "ATLAS instance started showing the bulkInsert constriant violation again. This is still occurring and we and the CERN developers intend to take advantage of this situation hopefully to finally identify the source of the bug." Sites need to absolutely make sure they have stable storage, tokens defined, solid CEs and batch systems. Then they are ready for data taking. ATLAS have been discussing other T1-UK T2 associations. - Other -- Biomed job mapping. Which sites saw biomed jobs running under another VO? (or any other similar mismatch?)
    • 11:25 11:35
      ROC stuff 10m
      ROC update *************** Last week saw the release of gLite3.1 Update28 The release contained * glite-CONDOR_utils for lcg-CE(PATCH:1856) * New version of gsoap plugin with a vulnerability fix (affecting LB, WMS, UI, WN, VOBOX, CE)(PATCH:1846) * Several bug fixes on WMS and clients (PATCH:1780) * New Short Lived Credential Service (SLCS), allowing to get short-lived personal certificate based on Shibboleth AAI identity (PATCH:1693) * MyProxy? version 1.6.1-7 (fixes build issue related to globus flavour, already deployed in production) (PATCH:1978) * Various improvements on lcg-extra-jobmanagers (CE) (PATCH:1942) * GFAL and lcg_util update with new function gfal_removedir and Several bug fixes * FTS SL4 release (32 and 64 bit) This version has a critical bug and should not be installed. The RPMs have been removed from the repository. This situation arose as the developer spotted the problem at the time of release - so it had passed certification tests. - Raised some TB-SUPPORT questions on automated updating -- How many sites were impacted? - One issue from last month is CA updates and checks (SAM vs repo update) - Sites in downtime for more than a month are now automatically suspended - Please add purchase information to the wiki: http://www.gridpp.ac.uk/wiki/Guidance_and_recent_purchases. This helps capture and share current thinking on hardware WLCG update ***************** FTS SL4 - required by the experiments or tier-1 sites? Alice: Neutral (as long as there is no disruption to the service. ATLAS: Prefer not to; to avoid introducing problems this close to data taking. CMS: Priority is stability for data taking days. Whatever is scheduled in advance *and* allows some pre-testing can be negotiated, though. On CERN migration, instead, PhEDEx /Prod vs /Debug instance can be played with to allow testing before going into prod (talked to Gavin) LHCb: Neutral (as long as there is no disruption to the service. There is an ATLAS jamboree this week: http://indico.cern.ch/conferenceDisplay.py?confId=38738 There is a GDB in September: http://indico.cern.ch/conferenceDisplay.py?confId=20233. This is now on 9th September. It was to look at T2 issues but that meeting is now to be held in October.
    • 11:35 11:45
      Security - news and discussion 10m
      - Follow up on recent incident - General discussion on incident response
    • 11:45 11:55
      Storage 10m
      https://twiki.cern.ch/twiki/bin/view/LCG/GSSDCCRCBaseVersions
    • 11:55 12:00
      AOB 5m