Operations team & Sites minutes Tuesday, 4 February 2014
Apologies: Wahid B, Andy W
Present:
Brian Davies (BD)
Alessandra Forti
Andrew McNab
Daniel Traynor
Daniela Bauer
David Crook
Duncan Rand
Elena Korolkova
Ewan McMAhon
Ewan steele
Gang Qin
Gareth Roy
Gareth Smith
Janu
Jeremy Coles
John Bland
John Hill
Matt Doidge
Matt Raso-barnett
Mohamid Kashif
Pete Gronbech
QMUL
Raja Nandakumar
Robert Frank
Sam Skipsey
Steve Jones
--------------------------------------------------------------
Experiment problems/issues (20')
- LHCb
(RN)
MC simulation
ARC CE support issue
Bristol, RALPP problem ruinning jobs
Problems submitting jobs to ARC CE
- CMS
Nothing Urgent
- ATLAS
http://indico.cern.ch/getFile.py/access?contribId=0&resId=0&materialId=slides&confId=291925
(JC) When do sites need to provide Multicore queues.
Needed for Run2
Users still using single core
Production is moving to Multicore
Request clarification of ATLAS request.
IS accounting accounted for?
what about inefficiency due to empty slots?
- Other
Enabled proxy-renewal on WMS at IC. Others are out of date
T2K going to storage meeting tommorow.
-------------------------------------------------------------
With reference to: http://www.gridpp.ac.uk/wiki/Operations_Bulletin_Latest
--------------------
Task Areas
General updates
Tuesday 4th February
The agenda for the February GDB is available.
The March pre-GDB will be on batch systems.
Discussion on RIPE ATLAS probes has continued off list. The PMB agree that there is an opportunity here and prefer to link this with outreach and dissemination. For those interested a discussion of what to propose will take place this Friday 7th February (email Jeremy).
--------------------
WLCG Operations Coordination - Agendas
Tuesday 4th February
There is a multi-core TF meeting this afternoon. The focus is on CMS and PIC.
The second middleware readiness meeting takes place this Thursday 6th February.
There was a ops coordination meeting last Thursday. The minutes are available. In summary:
BASELINES: WMS baseline downgraded to 3.6.1 for issues; APEL baselines added after meeting
OpenSSL: WMS needs new version of glite-px-proxyrenewal. ETA this week.
SAM: plan to split SAM services for WLCG (at CERN) and EGI (at consortium). Code will fork.
ALICE: gearing up for Quark Matter 2014 (May 19-24, GSI Darmstadt)
ATLAS: Rucio renaming campaign almost over. Rucio commissioning has started. DC14 simulation started on 1st of January.
CMS: DBS migration has been postponed. gLexec test (not yet critical) is a bit difficult for Tier-1s.
LHCb: Issues with ARC CEs.
FTS3: Experiments re-started increasing the load on the RAL FTS3 instance. Deployment discussion in February meeting.
gLexec: 22 tickets remain open. EMI gLExec probe (in use since SAM Update 22) crashes on sites that use the tarball WN.
IPv6: Report at next meeting.
MW readiness: Next meeting on Feb 6 at 15:30 CET: agenda in particular about how to involve experiments and sites. Need site input on table.
MULTICORE: October 2014 proposed by TF coordinators as a target date for a functional system to be deployed,
perfSONAR: New release this week. Lots of minor fixes and improvements. All sites should update to this.
SHA-2: EOS SRM for LHCb not yet OK. voms-proxy-init on lxplus crashes on creating SHA-2 RFC proxies.
TRACKING: No update
WMS decom: Deadline - end of April to decommission CMS and shared instances
--------------------
Tier-1 - Status Page
Tuesday 4th February
CVMFS Client version 2.1.17 has been rolled out on one batch of worker nodes (around 10% of the farm). So far so good.
Work going ahead preparing for various changes to the Tier1 network.
--------------------
Storage & Data Management - Agendas/Minutes
--------------------
Accounting - UK Grid Metrics HEPSPEC06 Atlas Dashboard HS06
Tuesday 4th February
A review of the HEPSPEC page shows no SL6 (or equivalent) entry for: UCL; Lancaster; Liverpool; Durham; ECDF; Glasgow; Birmingham; RALPP and RAL Tier-1.
The accounting pages show the following sites as not up-to-date with publishing accounting data: Lancaster (minor); RALPP and Sussex.
-------------------
Documentation - KeyDocs
See the worst KeyDocs list for documents needing review now and the names of the responsible people.
--------------------
Interoperation - EGI ops agendas
Tuesday 4th February
Short meeting yesterday, Agenda: https://wiki.egi.eu/wiki/Agenda-03-02-2014
SR: mpi v. 1.5.3
lb v. 4.0.12
apel-parser v. 2.2.1 and apel-ssm v. 2.1.1
Globus 5.2.5:
gridftp v. 5.2.5
gram5 v. 5.2.5
Still open WMS issues
EMI-2 decommissioning deadlines: 30/04/14 end of support, 31/05/14 deadline for upgrades
Affected:
ARC v2.*
ARGUS v1.5.*
BDII Site older than v1.2.0
BDII Top older than v1.1.0
CREAM v1.14.*
dCache v2.2.*
DPM older than v1.8.6
EMI-UI v2.*
EMI-WN v2.*
FTS v.2.2.8
StoRM older than v.1.11.0
VOMS v.2.*
--------------------
Monitoring - Links MyWLCG
--------------------
On Duty
Monday 3rd February
Good week. APEL ticket about Brunel alarms still open (although was passing this afternoon)
--------------------
Rollout Status WLCG Baseline
--------------------
Security - Incident Procedure Policies Rota
--------------------
Services - PerfSonar dashboard | GridPP VOMS
Tuesday 4th February
There is an update to perfSONAR (v3.3.2)
--------------------
Tickets
Monday 3rd February 2014, 14.30 GMT
Only 29 open tickets in the UK at the moment. To split it further, only 4 of these are "green", three are "yellow, the rest are "red". 7 are perfsonar related tickets, the only really big group of tickets we have.
RALPP
https://ggus.eu/ws/ticket_info.php?ticket=100480 (23/1)
Some obsolete entries were being published at RALPP, Chris thinks he has fixed it though (a problem on the cluster BDII), awaiting confirmation. Waiting for reply (31/1) Update-Solved
https://ggus.eu/ws/ticket_info.php?ticket=100849 (29/1)
Duncan has ticketed RALPP over their perfsonar latency box, he reckons a full log partition. Looks like this ticket hasn't been noticed yet though. Assigned (30/1)
OXFORD
https://ggus.eu/ws/ticket_info.php?ticket=99642 (10/12)
Backup Voms server testing for GridPP and Southgrid VOs at Oxford. On hold (30/1)
BRISTOL
https://ggus.eu/ws/ticket_info.php?ticket=99910 (20/12/2013)
LHCB having problems with the environment at Bristol, tracked to ARC being an odd duck. The problem has been forwarded to the ARC devs. On hold (21/1)
GLASGOW
https://ggus.eu/ws/ticket_info.php?ticket=98253 (21/10/2013)
Getting CMS working at Glasgow - the ticket. Gareth has updated a magic CMS xml file using one given to him by Daniela and notes that they're still failing CMS xrootd tests. Gareth asks if the tests are critical, and if they are he pleads for help. The lack of CMS credentials is really nobbling their efforts to getting this sorted, or even digging up docs. Waiting for reply (3/2) Update- Daniela provided an update containing what I can only assume is an invocation of dark forces, Gareth has risked his immortal soul and applied it.
EDINBURGH
I'll probably be better off coming back to these in a few weeks time!
https://ggus.eu/ws/ticket_info.php?ticket=100840 (29/1)
ECDF have an APEL-Pub nagios error going on. Looks like this has flown under the radar, probably due to both Andy and Wahid having more important things on their mind right now. Assigned (29/1)
https://ggus.eu/ws/ticket_info.php?ticket=99179 (25/11/2013)
Glue2 obsolete entries. Plans to retire the CEs have been slowed down due to waiting on networking changes. Andy reported that he'll fix the publishing if their not in position to decommission soon. On hold (24/1)
https://ggus.eu/ws/ticket_info.php?ticket=99180 (25/11/2013)
Similar to above, but publishing default values. It's the same CEs at fault, so this ticket is in the same boat. On hold (4/12/2013)
https://ggus.eu/ws/ticket_info.php?ticket=99794 (16/12/2013)
ECDF's perfsonar boxen blocking access to their webpages. Was held up by Christmas, but no news since-probably won't be for a few weeks. On hold (16/12/2013)
https://ggus.eu/ws/ticket_info.php?ticket=100569 (28/1)
The perfsonar latency box has started refusing connections. On hold whist Andy's off. On hold (28/1)
https://ggus.eu/ws/ticket_info.php?ticket=95303 (1/7/2013)
glexec ticket. Sadly the same story as last time (or the last times).
DURHAM
https://ggus.eu/ws/ticket_info.php?ticket=99621 (10/12/2013)
Durham have a bad worker node, spotted by enmr.eu. Whilst the guys haven't had a chance to fix it, one could argue that an offlined problem is a solved problem, as it can't hurt the jobs anymore. On hold (28/1)
SHEFFIELD
https://ggus.eu/ws/ticket_info.php?ticket=100037 (3/1)
Sheffield's perfsonar box needed some site firewall holes poking for it. On the to do list is an upgrade and assimilation into the mesh due to only testing against 6 sites currently. On hold (27/1)
MANCHESTER
https://ggus.eu/ws/ticket_info.php?ticket=100867 (30/1)
Teething problems for Manchester's new perfsonar boxes. Alessandra asks Duncan if it can be closed. In progress (3/2) Update- Solved, and wasn't a site problem to begin with.
LANCASTER
https://ggus.eu/ws/ticket_info.php?ticket=100566 (27/1)
Lancaster isn't getting 10G performance out of its perfsonar boxen. My suspicion is that the NICs themselves are running slow, not the switches. Maybe I'm using the wrong drivers? In progress (3/2)
https://ggus.eu/ws/ticket_info.php?ticket=95299 (1/7/2013)
Lancaster's GLEXEC ticket, waiting on me getting a tarball one working. I'm currently trying out another tarball one on my test bed, but it's early days yet (it's more an exercise in documenting the errors at the mo). On hold (31/1)
https://ggus.eu/ws/ticket_info.php?ticket=100011 (31/12/2013)
Biomed stopped working for one of the Lancaster CEs. The ticket suffered from lack of priority (sorry biomed!). On hold (24/1)
UCL
https://ggus.eu/ws/ticket_info.php?ticket=95298 (1/7/2013)
The UCL glexec ticket. SL6 and DPM upgrades are done, Ben is just getting things settled before he starts tackling this. On hold (27/1)
QMUL
https://ggus.eu/ws/ticket_info.php?ticket=94746 (10/6/2013)
QM having trouble scrubbing the biomed out of their SE's information system. Chris submitted https://ggus.eu/ws/ticket_info.php?ticket=100290 and has put a lot of hours into this. On hold (14/1)
BRUNEL
https://ggus.eu/ws/ticket_info.php?ticket=100568 (28/1)
Brunel's perfsonar have problems. Raul plans to upgrade, and has let know his distaste that an upgrade requires a reinstall. In progress (29/1)
EFDA-JET
https://ggus.eu/ws/ticket_info.php?ticket=97485 (21/9/2013)
LHCB job problems still haunting jet. I think this ticket should be in "Waiting for reply", but I also think that I know the answer to the question (that the error message they're seeing as a red herring). In progress, should be in some other status (29/1)
TIER 1
https://ggus.eu/ws/ticket_info.php?ticket=100114 (8/1)
Chis has spotted jobs failing to get from RAL WMS to Imperial. Looked to be SSL problems. On hold awaiting RAL upgrade to the next WMS release. On hold (30/1)
https://ggus.eu/ws/ticket_info.php?ticket=100343 (16/1)
RAL WMS producing 512-bit proxies (occasionally). Waiting on the same release. Waiting for reply (?) (27/1)
https://ggus.eu/ws/ticket_info.php?ticket=100887 (31/1/2013)
Due to the same underlying issue as the above tickets , Chris asks for the gridsite package on the webdav LFC to be updated. In progress (31/1)
https://ggus.eu/ws/ticket_info.php?ticket=100507 (23/1)
CMS transfers failed between Caltech and RAL. The problem has eased itself, so the ticket only needs to be kept open if further investigation is warranted (as Brian pointed out). In progress (3/2)
https://ggus.eu/ws/ticket_info.php?ticket=98249 (21/10/2013)
CVMFS for SNO+. Almost there, creating the Sno+ tarballs to test with is taking longer then expected. On hold (29/1)
https://ggus.eu/ws/ticket_info.php?ticket=99556 (6/12/2013)
The new NGI Argus server (argusngi.gridpp.rl.ac.uk) has been set up in the gocdb and is online. In progress (30/1)
https://ggus.eu/ws/ticket_info.php?ticket=97025 (3/9/2013)
Ye olde RAL myproxy server name confusion issue. No news on this for a while, the hope is having this dealt with soon. But then the last update was nearly a month ago, so soon isn't as soon as we'd like it to be! On hold (6/1)
--------------------
Tools - MyEGI Nagios
--------------------
VOs - GridPP VOMS VO IDs Approved VO table
--------------------
Site Updates
--------------------------------------------------------------------
EGI OMB news/updates (15')
With reference to the OMB on 30th January: https://indico.egi.eu/indico/conferenceDisplay.py?confId=1858
- UMD-2 will be decommissioned in the coming months. 30th April end of security support. 31st May all services to have been removed or upgraded.
- UK sites failing Glue2:
RAL-LCG2
UKI-LT2-IC-HEP
UKI-NORTHGRID-MAN-HEP
UKI-NORTHGRID-SHEF-HEP
UKI-SCOTGRID-GLASGOW
UKI-SOUTHGRID-RALPP
Check with the glue2 validator to see errors.
- A new Operations Dashboard will be in pre-production during February and moved into production in March.
- Availability/reliability targets for EGI are moving to 80%/85%.
- Some new ARC SAM tests are being introduced: org.nordugrid.ARC-CE-LFC-result; org.nordugrid.ARC-CE-LFC-submit; org.nordugrid.ARC-CE-SRM-result;
org.nordugrid.ARC-CE-SRM-submit; org.nordugrid.ARC-CE-submit
- There is a summary page on how to publish from various middleware types: https://wiki.egi.eu/wiki/MAN09
- There was an overview of the French NGI adoption of iRODs.
- FedCloud to production: Management (GOCDB - ok); Monitoring (in progress); Accounting (in progress); Documentation (ok); Support (ok); Dashboard (in progress) and Security (in progress).
- The SAM tests are: org.nagios.CloudBDII-Check; eu.egi.cloud.OCCI-VM and
org.nagios.OCCI-TCP. Security checks not yet available.
--------------------------------------------------------------------
Status checks (10')
- VOMS network enablement
https://www.gridpp.ac.uk/wiki/Adoption_of_Backup_GridPP_Voms_Servers
Six sites done
--------------------------------------------------------------------
AOB (1')
- Dissemination
http://planet.gridpp.ac.uk/
- Suggestions for GridPP32 (registration now open http://www.gridpp.ac.uk/gridpp32/).
Chat Window:
[11:10:39] Jeremy Coles In the ATLAS update "sites should provide multicore resources.. dynamic preferred... no pressure on sites to have multicore queue before April"
[11:15:34] Elena Korolkova I'll ask for official statement on multicore request
[11:16:14] Jeremy Coles Thanks.
[11:23:39] Jeremy Coles https://www.gridpp.ac.uk/wiki/Operations_Bulletin_Latest#
[11:25:34] Duncan Rand what's the odd noise?
[11:26:08] Alessandra Forti someone hasn't muted
[11:26:28] Duncan Rand QMUL?
[11:27:11] John Bland brian
[11:27:16] John Bland do you need to mute?
[11:30:03] Ewan Mac Mahon I'm almost inclined to just close is as not solved.
[11:30:15] Ewan Mac Mahon I think it's passed the point of relevance.
[11:34:14] Elena Korolkova Indeed, thanks to John Bland
[11:37:09] Ewan Mac Mahon What was the symptom, and what's plugged into the card? Is it a twinax copper cable or an actual optic?
[11:45:13] John Hill I didn't realise it was March already - yikes!
[11:47:29] Matt Doidge Thanks Ewan - the symptom is the NICs only seem to be running at 1G, whereas we don't seem to have any 1G bottlenecks in the network. They're both fitted with optics.
[11:48:04] Alessandra Forti sorry I dropped
[11:53:01] Matt Raso-Barnett Im here, but I have no mic I'm afraid. I'm looking into Sussex APEL issue but no progress so far
[11:54:46] Jeremy Coles Thanks
[11:56:09] Steve Jones It's on my worklist
[11:56:43] Ewan Mac Mahon I still want to give it a go, it's entirely a round tuit problem.
[11:58:42] Ewan Mac Mahon @Matt optics can be a bit funny in the X520 cards; there's some diagnostic stuff we can try but it's probably simpler by email.
[11:59:08] Ewan Mac Mahon Do you have any twinax copper cables you could try though, or are the links just too long?
[12:01:36] Ewan Steele hmm thats interesting I thought we were failing too
[12:04:31] Matt Doidge I'm afraid I don't have any twinax cables or the port-doohickies in the rack switch to take them. I'll poke you lovely chaps offline once I'm sure I'm not doing something stupid (tm)
[12:06:42] Jeremy Coles https://www.gridpp.ac.uk/wiki/Adoption_of_Backup_GridPP_Voms_Servers
[12:09:52] Jeremy Coles http://planet.gridpp.ac.uk/
[12:09:58] Steve Jones We've done the UIs and added the date to the table. Thanks for the reminder.
There are minutes attached to this event.
Show them.