IT Piquet Project meeting

Europe/Zurich
513-1-021 (CERN)

513-1-021

CERN

Veronique Lefebure (CERN)
Description
Modifications to the LCG MoU and Piquet Services organisation
Present: Maria Girone, Jamie Shiers, Nick Thackray, Ludwig Pregernig, Veronique Lefebure LCG MoU Jamie provided an example of service recovery taken from the CMS CSA06 experience: service was interrupted at two occasions. In one case, a DB problem was fixed by the DB piquet (reboot needed), in the other case there was a netwrok problem caused by a human intervention (during day time). In both cases, CMS was able to increase their process rate and catch up. Some suggestions for modifications will be made regarding the LCG MoU content:
  • There are a number of inconsistencies in the MoU tables that need to be addressed. For example, the targets at both T0 and T1 of the data export end need to be symmetric.
  • There is also the loop-hole of scheduled downtime. It is not acceptable for a site to declare downtime for the next two decades and therefore get off the hook. However, the MoU currently allows this...
  • Secondly, the wording is somewhat mainframe-era-ish. It is not realistic to pretend that we have mainframe-style operation. A more realistic analogy would be with a reliable parcel delivery service, such as FedEx. (To ask if FedEx is up all the time is entirely meaningless. However, the inability to send a parcel, track it, or its non-delivery are meaningful service measures. Machine operation plans Maria:
  • both during end of 2007 (nov & Dec) and 2008 machine operations are not considered as "normal": - end of 2007: beam during Nov, Dec - shutdown during Jan, Feb, Mars 2008 - beam starting in ~April 2008 (commissionning, then physics): machine developmentd during the day, fill in the evening, stable beam over night ===> data taking over night
  • operations during 2009 and later are planned to be as follows: -140 to 160 days of physics per year (winter excluded) -each month: ~20 days of physics, followed by ~3 days of machine development and 3-4 days of no beam. ===> "smooth" running needed 3 weeks out of 4 ===> interventions possible or to be planned only once per month Jamie: - Special Piquet will be needed both end of 2007 and 2008 because of tension, and service immaturity - The period during which Special Piquets are needed should not be too long: better keep the money to hire a developper and make the software robust (for CASTOR at least) But: for DB services, there will always be the need of a Piquet (Oracle support can not be done by operators/sysadmins alone). For FTS: the recent experience showed that a simple restart of the services has always been the solution to the problems, and this can be given to the the operator/sysadmin piquets. Jamie reminded that, for data transfer, piquet on both sides are needed. Nick reminded that Remedy is an important element of the GRID problem tracking, as some alarms are directly fed to Remedy for contacting responsible persons.
  • There are minutes attached to this event. Show them.
      • 15:00 15:40
        LCG MoU revision 40m
        In the LCG MoU document of October 2006, it is said that "the minimum levels of service will be reviewed by the operational boards of the WLCG Collaboration". This working group is the good moment to at least initiate the process.
        Speaker: Dr Jamie Shiers (CERN)
        Service Recovery illustration
      • 15:40 15:55
        Machine Operations Plan 15m
        Three possible scenario for 2007/2008
        Speaker: Maria Girone (CERN)
        Maria Slides
      • 15:55 16:10
        Service Review 15m
        Continue Service review (dependencies, criticality, Piquet coverage)
        New diagram
        Piquet availability
      • 16:10 16:15
        Next meeting 5m
        Date for next meeting + Who shall we invite ?