Currently over 80 sites are connected to the LCG grid and over 8000 processors are available to run a variety of applications. During the Data Challenges of the LHC experiments the grid middleware has proven to be stable though incomplete. It seems a good moment to shift the attention from getting more reliable software to getting more reliable grid operation. In the end we need an infrastructure which is always operational, where software upgrades can be done regularly and in a controlled way, where bugs can be fixed quickly and efficiently and where users can get support when and where needed. To discuss how to achieve this a Workshop will be organised at CERN from 2 to 4 November. We would like to make this a real workshop with one plenary sessions only followed by many small dedicated meetings focused on just one aspect.
For this open workshop people responsible for the operation of the major LCG centers are invited as well as the people responsible for the EGEE Operational Management Center, the Regional Operation Centers and Core Infrastructure Centers. The people with the real hands-on experience of operations should come to propose solutions for the bottlenecks we will identify as well as the managers that can assign resources and manpower to make these solutions become true.
The format of the workshop is outlined below. The IT Auditorium can hold approximately 100 people and we would therefore like to ask you to register by filling up the form at the following web address:
http://lcg.web.cern.ch/LCG/SC2/LCGWorkshop/LCGWorshopReg.asp
Although this is an open workshop the organizers retain the right to make some choices in case of over subscription.
Summary of OSG/EGEE/LCG incident response procedures and plans + discussion
Speaker:
Dave Kelsey(RAL)
transparencies
4
Overview of the deployment process
Speaker:
Markus Schulz(CERN (IT-GD))
transparencies
5
Operations management
Overview of site testing, problem follow up and escalation. Local vs remote control of sites.
Discussion on escalation procedures. How to handle bad sites?
Speaker:
Piotr Nyczyk(CERN (IT-GD))
transparencies
12:30
LUNCH
6
Grid3 Operations
Speaker:
Doug Pearson
transparencies
7
Sharing work and responsibilities
Summary of what was discussed at CHEP. How to share the work between the CICs and other GOCs.
Speaker:
John Gordon(RAL)
transparencies
8
Globus Monitoring
Speaker:
Jennifer Schopf(ANL)
transparencies
9
Monitoring Frameworks
Overview of R-GMA monitoring framework. What information needs to be published by each site?
The working groups report on what they have achieved. The slides should show an operational plan with a schedule when this can achieved and with names of people who are responsible for it.