Attendance list:
Jeremy Coles (Chair)
Sam Skipsey (Minutes)
Daniela Bauer
John Bland
Chris Brew
Stephen Burke
James Cullen
Chris Curtis
Santanu Das
Brian Davies
Matt Doidge
Alessandra Forti
Pete Gronbech
Rob Harper
Stephen Jones
Mohammad kashif
Elena Korolkova
Raul Lopes
Winnie Lacesso
Ewan Mac Mahon
Dug McNab
Raja Nandakumar
Stuart Purdie
Duncan Rand
Derek Ross
Govind Songara
Stuart Wakefield
Chris Walker
Apologies: Wahid Bhimji, Graeme Stewart
===Minutes begin:
==ROC update:
Dug: nothing to report from on-duty. No ticketing issues.
It was suggested that we should look back at history in June to compare reliability over time.
==EGEE Ops meeting:
=gLite 3.2 update. Torque release for SL5. SCAS/gLexec/CREAM out. (This is significant because it means there's a stable release of SCAS and gLexec for SL5 now.)
=gLite 3.1 update 60 - also has the new Torque. WMS + CREAM updates. No mention of staged-rollout issues.
=Early Adopter sites now exist (the successor of PPS) - UK would like to get more involved.
*Can sites interested talk to Jeremy?
==T1 news:
Gareth has blogged about major interventions due next week: draining batch system, fscking disks, etc. This will happen around the 27th-28th of Jan, and most services will be affected.
UPS bypass test showed that the UPS *is* the problem for current noise. Workarounds are being put in place so databases can move back.
glexec/SCAS being worked on by Derek.
Frontier servers will also be upgraded soon.
==WLCG Update:
Graeme was our man at the GDB this time. He has filled out a summary report on it (linked in agenda).
The following items were mentioned:
= GGUS availability: had some issues with connections to GOC-DB and site name consistency. Fixed. Some discussion of frequency of testing T1 alarm systems.
= Maarten Litmaath's discussion of issues brought up by the wlcg tech forum. First issue is MUPJs (multi-user pilot jobs). The forum have decided to get more site feedback from questionnaire - what does a site do if job switching doesn't work, etc?
= Updates on Joint Security Policy Group policies. Comments by the end of Jan. (See Dave Kelsey's talk. Mingchao will push around an email.)
= Middleware (Oliver Keeble): proposal to drop releases after 2 months of stable running of next release (i.e. 3.1 -> 3.2), "stable" as defined by weekly operations group.
Brian: I understood that 3.1 is the SL4 tree. Are we now saying that within two months people should upgrade to SL5 for 3.2 once a 3.2 stable version works?
Jeremy: yes, but all that happens is that the 3.1 version is not supported - not forced to upgrade.
Jeremy will email TB-SUPPORT with mention of this.
Oliver Keeble also mentioned versioning - request for feedback on a proposal about rationalising metapackage versions to be meaningful to site admins.
= OSG activities update.
Experiment priorities - ALICE prioritise SL5. Have tested CREAM. ATLAS most concerned with CASTOR stability at CERN, also SRM monitoring. CMS - happy, some discussion about cache space on WNs (20Gb/core). LHCb - no severe problems, but still need to download data to WNs on dCache sites due to dCap issues.
==Ticket statuses:
=Chris Walker confirms that Bristol's STORM publishing issue would be fixed by moving from 1.3 -> 1.4. 1.5 is due soonish, if Bristol want to consider it?
=Close fusion tickets due to lack of response.
=Camont issue at Bham: Chris Curtis is looking at it.
==Expt problems and Issues:
= LHCb - Raja.
Nothing much to say this week. Just mostly running user jobs. See a couple of tiny problems in WNs, but nothing major.
SAM tests, especially towards T1 SRM are flaky and not reliable - being worked on now.
= CMS - Stuart Wakefield via phone.
Not much to say. Some small issues with software areas - QM's software area is too large (>2Tb size breaks their apt-get installs!, needs workarounds. This is relevant because Lustre is better than NFS for reliability, but the large filesystem sizes are problematic for old utilities.), Imperial software area for SL5 fell over under load, being ported.
One small issue with ECDF and transfers.
= ATLAS - no report. No particular problems known of.
= Other -
CAMONT - Mark Slater set up tracking at Bham for jobs they're running. Consider ones that fail or are killed - problem with Oxford (apparently because it's going to a CREAM queue - this may be an issue at CREAM CEs in general).
Kashif: it was a problem with the CREAM CE's queues being closed, it's a small batch system with not many slots. (The 600 "killed" jobs currently waiting were just being queued.)
Is CAMONT targeting queues unfairly?
Why does svr014 at Glasgow also have a high kill rate?
Dug: We don't know! We shall investigate - we haven't been complained at directly about failure rates, so we didn't know there was a problem!
It might be fairshare related - CAMONT can only run a maximum of 100 jobs at Glasgow, so… (if jobs are being killed due to being queued a long time, then…)
Kashif: Mark says he kills the jobs if they're queued for a long time. So, if there are a lot of jobs that turn up and can't run due to limits…
Currently: problems at Cambridge and Glasgow. Check if we see any jobs running at all.
QMUL is trying to support CAMONT but isn't on the list!
Fusion VO: Matt has problems at Lancaster getting authorisation to work.
Daniela has fixed this for Matt with the magic power of fresh eyes.
==Site issues:
= Oxford:
Pete: aircon is completely repaired and all is working. Need to review to make sure it is all working etc.
Kashif: SCAS is working in production mode, on a different limited batchsystem with only 8 job slots against a CREAM CE. No glexec installed on WNs yet. Probably next week.
= Bristol: Winnie:
Some instabilities with VM CEs; rebuild StoRM 1.3=>1.4; new VMs SL5 site-bdii & SL4 MON
= Liverpool: Jon: Had some problems with accounting to APEL. Looks like it is publishing now, but portal hasn't updated for some time. Working on a VM CREAM CE for testing. Thinking about where and how to spend procurement money - link with central cluster doesn't help, as they're slow to upgrade.
= T1: Derek: SCAS/glexec looked at in the last couple of days.
= Lancs: Matt: Just got a new server rack. Second CE replacing old CE, new CREAM CE as VMs.
= RALPP: Chris Brew: In downtime this week due to power work last weekend this weekend - taking the chance to make big changes, splitting dCache headnode to move SRM off (gone well, probably), redoing all pool accounts to get a unified UID/GID space that is sensible, splitting off primary group of local users in VO from primary groups of pool accounts (T3/T2 scheduling). A lot of updates, then. Hoping to stick an SCAS somewhere, too - just need to get hardware for it.
Pete: slight issue with RALPP accounting - it seems to be underpublishing by factor of 2 - 3.
Chris: yes, when we moved to HEPSPEC06, APEL hasn't coped with the subcluster setup, and is still using the old number. Ticket open with APEL to resolve this, once we come out of downtime.
Pete: this might have been a problem with Cambridge, also.
= Bham: Chris: Been focussing on SL5 on shared cluster. Basic stuff works, but some trouble with ATLAS software (will be working in a week, completely SL5).
= Glasgow: Dug: Waiting on ATLAS to test GLEXEC/SCAS, rolling update of wns, commissioning VM servers, investigating pilot factory/cream issues.
= Sheffield: Elena: as a major issue we are waiting for money to upgrade the cluster, Set up additional storage. The problem is to publish storage in Gstat 2.0. I think we should reyaim DPM
= Imperial: Daniela: Gremlins in the machine room. Tried to set up new software area, but machine failed for no good reason (drafted in WN as temporary measure), problems with storage nodes (investigating these). (Upgrading dCache to Chimera not that high on priority list.)
= Brunel: Duncan: More storage to be installed to 500T! New CE in new machine room to be assigned. Experiment software server needed. Retiring old CEs, moving to the new room. Infrastructure work. All to do with new machine room.
= Cambridge: Santanu: Preparing for SL5 migration - waiting on new expt software area location. 2Tb disk being set up for it.
Had a problem with APEL which is being looked at.
= RHUL: Govind: Next week, downtime to move from Imperial to new machine room. At the same time will migrate to SL5. 3-4 week downtime expected. Have informed all the VOs so they can move data etc if needed.
Duncan: moved the monte carlo back to Imperial so it can continue going.
= QMUL: Chris: Main push is to SL5. Have a system testing with 40 machines, but some kept dropping Lustre - and this broke the metadata server! Solving this problem is the main issue. Looking to procure new hardware - going to tender in the near future.
StoRM 1.5 upgrade due when it comes out, 10GigE for the storage head node (as it got DDoSd by h1 accidentally!).
(The problem is that the jobs get data by gridftp, and thus load the StoRM node, not the Lustre filesystem with the local transfer protocol!)
= Manchester: James: Working on glexec on WNs/SCAS. Had it sort of working on Friday. Working on it more this week.
Had some new network cards delivered - bonding + checkingperformance in storage servers (HC test)
Since mid last week been suffering with problems with kickstarting new WNs. Not sure what the cause is.
= UCL-HEP: Gianfranco Sciacca: stable, running ATLAS production. New CE for SLC5 nodes. One SLC5 node online, got ATLAS sw installed over holidays on separate SW area. Next: move all WNs to SLC5. New storage almost ready to go online (50+ TB). Behind on squid, working on it now
= UCL-CENTRAL: Gianfranco Sciacca: stable over holidays, just got SAM error yesterday, didn't hear from admins about this yet. Ran ATLAS prod with good share of nodes. Ready to move UCL-CENTRAL WNs under UCL-HEP site-bdii
Brian: What is the timescale for the upgrade for dCache at Imperial? (Either to fastPNFS, or …)
Duncan: haven't finalised a time for this to happen. This isn't impacting CMS, so it is of lower priority than it might be.
Brian: The main issue is that Imperial's current version is becoming increasingly less supported by dCache developers.
(Covered SCAS/glexec section of meeting by the above reports)
== MUPJS:
Read page and fill in questionnaire please! Waiting on most sites for the response.
Maarten would like responses by this Friday! (Email Jeremy before the deadline so he can compile the response.)
(John Gordon noted that Sites may wish to talk to their security people and managers before filling it in.)
( Mingchao's Security questionnaire: QMUL, RAL T1 haven't responded yet?
QMUL response will happen when SL5 is up and discussion has been had. )
== gStat2 publishing issues:
Started ticketing sites with issues.
= Follow the link in the agenda.
If your site isn't in the list, then you're not visible as being part of GridPP.
A number of issues with Storage and Waiting Jobs publication.
As an example, Brunel: Duncan was discussing the physical/logical cores publishing with Raul, seems okay.
Published used online space is okay, too.
[There was some time spent trying to properly interpret the meaning of the values on the page itself!
Stephen Burke: we need to know what the underlying algorithm is to understand the Total values.
Chris: QM double counts the jobs as it has two CEs in front of one cluster.
John Gordon: this isn't productive without feedback from Steve Traylen!]
There was some discussion on TB-SUPPORT about how to correctly publish CE values.
We don't understand the job Waiting figures.
Cambridge Storage publishing is obviously wrong. (=0)
Santanu will look into it. XML file was pointing at the wrong node for a while.
John Gordon: In gstat, you can look at the ldap view and see which BDIIs it is looking at.
Chris Brew: we fixed a number of bugs with Steve Traylen before Christmas, so can offer advice. Need to be pretty accurate about your scaling factor if you're using one - the result must be within 1 of the logical CPUs!
Should ask Steve T for pointers.
Jeremy will start ticketing the obvious problems.
Jeremy will also cross-check with quarterly reports.
==Actions:
=Continue to update resiliency page on the wiki.
=Continue to update the hardware page with procurement decisions.
==AOBs -
=1/4rly reports are due!
Pete has started his. Duncan has a draft already.
= Stephen Burke: it looks like ATLAS really are transitioning to SL5 (only) native soon! A reason to move when you can!
Chris+Jeremy+Raja: as are CMS and LHCb.
Meeting ends.
Chat window log:
[10:59:05] Mohammad kashif joined
[10:59:27] Winnie Lacesso joined
[10:59:36] Stephen Burke joined
[11:00:04] Jeremy Coles joined
[11:00:20] Stephen Jones joined
[11:00:40] raul lopes joined
[11:00:49] Chris Curtis joined
[11:01:28] Dug McNab joined
[11:01:55] Ewan Mac Mahon joined
[11:02:00] Ewan Mac Mahon left
[11:02:14] Pete Gronbech Hi All, I've forgotten my headset so will avoid speaking but can hear ok.
[11:02:19] Raja Nandakumar joined
[11:02:24] Ewan Mac Mahon joined
[11:02:37] Elena Korolkova joined
[11:02:38] Duncan Rand joined
[11:03:20] Chris Brew joined
[11:03:44] Daniela Bauer joined
[11:03:55] Stuart Purdie joined
[11:06:52] Stuart Wakefield joined
[11:07:12] Jeremy Coles We started with the ROC update and will return to Expt problems and issues.
[11:07:47] Rob Harper joined
[11:10:22] Santanu Das joined
[11:12:11] Govind Songara joined
[11:12:38] Queen Mary, U London London, U.K. joined
[11:14:28] James Cullen joined
[11:15:03] Duncan Rand apparently the phone bridge doesn't work
[11:15:19] Duncan Rand code's invalid
[11:18:23] Phone Bridge joined
[11:18:48] Stuart Wakefield left
[11:18:55] Winnie Lacesso Need to rebuild StoRM 1.3 to 1.4 (no upgrade possible) It publishes in bytes
[11:20:03] Winnie Lacesso I've asked, the stoRM developers say no upgrade possible.
[11:22:29] Gianfranco Sciacca joined
[11:23:34] Alessandra Forti I have to go
[11:27:24] Pete Gronbech your sound just dropped out and then returned [11:27:50] Alessandra Forti left
[11:28:10] Dug McNab I believe Graeme is off ill at the moment.
[11:30:32] Dug McNab Jeremy what are you looking at?
[11:31:06] John Gordon joined
[11:32:12] Jeremy Coles http://epweb2.ph.bham.ac.uk/user/slater/camont/currtest/ce_info.html
[11:35:17] Phone Bridge left
[11:35:18] Queen Mary, U London London, U.K. It is QMUL's intention to be supporting camont - but it isn't on the list above
[11:36:31] John Gordon left [11:37:29] Phone Bridge joined
[11:39:10] Winnie Lacesso Some instabilities with VM CEs; rebuild StoRM 1.3=>1.4; new VMs SL5 site-bdii & SL4 MON.
[11:46:42] Pete Gronbech I was just saying the Cambridge have the oposite problem their setup failed to pick up the new lower number on 1st oct so are still over publishing.
[11:47:46] Jeremy Coles Good to know for the Q409 report.
[11:48:19] Elena Korolkova as a majot issue we are waiting for money to upgrade the cluster, Set up additional storage. The problem is to publish storage in Gstat 2.0. I think we should reyaim DPM
[11:48:55] John Gordon joined
[11:49:38] Santanu Das I'm looking in the APEL problem, but din't able to figure out any thing yet
[11:53:47] Elena Korolkova Most of sites are publishing the wrong number of jobs running/ waiting. Any recipe how to fix this?
[11:54:24] Phone Bridge left
[11:55:36] Phone Bridge joined
[11:56:37] Gianfranco Sciacca UCL-HEP: stable, running ATLAS production. New CE for SLC5 nodes. One SLC5 node online, got ATLAS sw installed over holidays on separate SW area. Next: move all WNs to SLC5. New storage almost ready to go online (50+ TB). Behind on squid, working on it now
[11:57:51] Gianfranco Sciacca UCL-CENTRAL: stable over holidays, just got SAM error yesterday, didn't hear from admins about this yet. Ran ATLAS prod with good share of nodes. Ready to move UCL-VCENTRAL WNs under UCL-HEP site-bdii
[12:02:48] Chris Brew you just had RALPPs
[12:05:43] Derek Ross I've pushed the MUPJ questionaire round the team for feedback
[12:07:10] raul lopes yes it is
[12:08:51] Chris Brew RALPP is missing but I suspect that is because our site-bdii has been off since Friday
[12:11:26] Elena Korolkova JFor Sheffield, the number od CPUs yesterday was 200,both physical and logical. Why it is 0 today, I don't know. I'm wondering is there a link how to fix problem with publishing in gstat2.0
[12:11:27] raul lopes I can speak from here. Storage numbers are ok, but jobs running is incorrect.
[12:11:46] Elena Korolkova yes
[12:12:46] Phone Bridge left
[12:14:46] Elena Korolkova For Sheffield now it is the true number of CPUs as well as for jobs
[12:15:11] Dug McNab any number for physical and logical where they are the same and the cluster has multi core machine is probably wrong.
There are minutes attached to this event.
Show them.