Present:
Maria Girone, Jamie Shiers, Nick Thackray, Ludwig Pregernig, Veronique Lefebure
LCG MoU
Jamie provided an example of service recovery taken from the CMS CSA06 experience: service was interrupted at two occasions. In one case, a DB problem was fixed by the DB piquet (reboot needed), in the other case there was a netwrok problem caused by a human intervention (during day time). In both cases, CMS was able to increase their process rate and catch up.
Some suggestions for modifications will be made regarding the LCG MoU content:
There are a number of inconsistencies in the MoU tables that need to be addressed. For example, the targets at both T0 and T1 of the data export end need to be symmetric.
There is also the loop-hole of scheduled downtime. It is not acceptable for a site to declare downtime for the next two decades and therefore get off the hook. However, the MoU currently allows this...
Secondly, the wording is somewhat mainframe-era-ish. It is not realistic to pretend that we have mainframe-style operation. A more realistic analogy would be with a reliable parcel delivery service, such as FedEx. (To ask if FedEx is up all the time is entirely meaningless. However, the inability to send a parcel, track it, or its non-delivery are meaningful service measures.
Machine operation plans
Maria:
both during end of 2007 (nov & Dec) and 2008 machine operations are not considered as "normal":
- end of 2007: beam during Nov, Dec
- shutdown during Jan, Feb, Mars 2008
- beam starting in ~April 2008 (commissionning, then physics): machine developmentd during the day, fill in the evening, stable beam over night
===> data taking over night
operations during 2009 and later are planned to be as follows:
-140 to 160 days of physics per year (winter excluded)
-each month: ~20 days of physics, followed by ~3 days of machine development and 3-4 days of no beam.
===> "smooth" running needed 3 weeks out of 4
===> interventions possible or to be planned only once per month
Jamie:
- Special Piquet will be needed both end of 2007 and 2008 because of tension, and service immaturity
- The period during which Special Piquets are needed should not be too long: better keep the money to hire a developper and make the software robust (for CASTOR at least)
But: for DB services, there will always be the need of a Piquet (Oracle support can not be done by operators/sysadmins alone). For FTS: the recent experience showed that a simple restart of the services has always been the solution to the problems, and this can be given to the the operator/sysadmin piquets.
Jamie reminded that, for data transfer, piquet on both sides are needed.
Nick reminded that Remedy is an important element of the GRID problem tracking, as some alarms are directly fed to Remedy for contacting responsible persons.
There are minutes attached to this event.
Show them.