WLCG-OSG-EGEE Operations meeting

Europe/Zurich
28-R-15 (CERN conferencing service (joining details below))

28-R-15

CERN conferencing service (joining details below)

Nicholas Thackray
Description
grid-operations-meeting@cern.ch
Weekly OSG, EGEE, WLCG infrastructure coordination meeting.
We discuss the weekly running of the production grid infrastructure based on weekly reports from the attendees. The reported issues are discussed, assigned to the relevant teams, followed up and escalated when needed. The meeting is also the forum for the sites to get a summary of the weekly WLCG activities and plans
Attendees:
  • OSG operations team
  • EGEE operations team
  • EGEE ROC managers
  • WLCG coordination representatives
  • WLCG Tier-1 representatives
  • other site representatives (optional)
  • GGUS representatives
  • VO representatives
  • To dial in to the conference:
    a. Dial +41227676000
    b. Enter access code 0157610

    OR click HERE

    NB: Reports were not received in advance of the meeting from:

  • ROCs:
  • Tier-1 sites: BNL, Triumf, INFN, NDGF
  • Tier-1 availability reports:
  • VOs:
  • list of actions
    Minutes
      • 16:00 16:05
        Feedback on last meeting's minutes 5m
        Minutes
      • 16:01 16:30
        EGEE Items 29m
        • <big> Grid-Operator-on-Duty handover </big>
          From ROC UK/I (backup: ROC AsiaPacific) to ROC CERN (backup: ROC Central Europe)

          NB: Please can the grid ops-on-duty teams submit their reports no later than 12:00 UTC (14:00 Swiss local time).

          Tickets:
          Backup team:
          New opened: 20
          2nd mail :28
          Quarantine :16
          Close tickets: 31
          Extend :8

          Issues:
          1. #30175 : site name is csTCDie has RB problem but there is no info on GOCDB
          2. ru-Moscow-SINP-LCG2 has many alarms but it is in SD.
          3. sometimes can not close the alarm that site is ok.
        • <big> PPS Report & Issues </big>
          PPS reports were not received from these ROCs: Italy

          This week, as the last one, ther will be no pre-production release. This is intended to let the service recover from the issues experienced with UPDATE29:

          Issues from EGEE ROCs:
          1. (ROC SouthWest Europe): Last PPS-update (29) was not very smooth. Sites found problems with the configuration and CODs opened GGUS tickets on them. Sites believe this is not a problem of the sites, but of the PPS update, which was in a not very stable state when given to the PPS.
            - Should PPS sites put themselves in Scheduled Downtime everytime they install an update?
            - Is there a way to put the "whole PPS" in Shed Downtime when a new update is being rolled out?
            The idea is to avoid CODs opening several tickets on sites, if the problems are in the update itself.


          Speaker: Nicholas Thackray (CERN)
        • <big> Status update on Classic SE </big>
          The classicSE is completely frozen. This leaves the questions:

          1. The things that are not in the classicSE that people want added? From Gavin's comments these are already in DPM so DPM is an acceptable upgrade.
          2. What are the things in the classicSE that are not available in DPM or dCache? The obvious answer is real posix mounting of the file system probably via NFS. Both DPM and dCache are looking at supporting NFSv4, there are comments from both of these that NFSv4 might be available by end of the year or sooner.
          3. So the final question is will the classicSE still be included in the upcoming gLite release? The answer is yes that the classicSE will remain in the gLite 3.1 release.

          Once there is a DPM with NFSv4 support then this will be re-evaluated.
        • <big> Update on progress towards SL4 support </big>
          Build Status:
          • All of SL4 32bit builds except APEL Publisher because of some problem with javac, which is under investigation. There is another problem with lcg-info-dynamic-dpm.
          • 64bit is lower priority than 32 bit and while it is being built routinely by ETICS there are compile problems with the build.
          • Current status available: https://grid-deployment.web.cern.ch/grid-deployment//cgi-bin/reports.cgi?action=package
          Install Status:
        • Installations are now happening of some SL4 32 bits node types, in particular some installs okay e.g UI, some fails with dependency problems CE. and some have not been tried e.g WMS. The priority node type in this configuration is the UI.

        • Testing and Certification Status:
        • To early to give any status.
        • Having said all this repositories do exist of SL4 32bit builds if any one wants to try and feedback is appreciated but not necessarily acted upon with urgency.
          There are no time scales yet for completion.
  • <big> EGEE issues coming from ROC reports </big>
    1. (ROC Central Europe): [INFORMATION] 2) Stress tests shown that turning indices on at slapd server makes the things worse at a higher load. Probably due to indices are rebuilt at each a few minutes. See here for details: http://wiki.grid.cyfronet.pl/CoreServices/SLC4BDII#head-97efd498b5ee9f7e9a63f94ff0eb86adee659530


    2. (ROC DECH): [INFORMATION] SRM/SE-host-cert-valid tests on invalid ports (those tests have now been made not critical again, but are still failing at FZK without real problem): https://gus.fzk.de/pages/ticket_details.php?ticket=16776


    3. (ROC Italy): The recent release process has caused several issues in terms of bugs and annoyances. Just as examples:
      - SGM poolaccount issue, also reported by SEE-ROC last week
      - LFC-DPM and glite-yaim issues added in the gLite Update 24 page in a second time, after the problems appeared in production.
      - VOBOX issue reported by Maarten L. on LCG-ROLLOUT last friday (why not broadcasted?)
      We have followed and agree about the thread on PPS ML, with subject "A pause in the PPS updates flow?". It seems that the release flow is too fast, not fully checked, especially within the PPS step: there's not enough time to properly check the release in PPS, and perhapse not enough coordination of the Experiments effort to test the PPS release. Do we really need so frequent updates in production? We think that a better release process could make it easier.


    4. (ROC Russia): LFC publish information via site GIIS (CE). Therefore, if CE is down, the corresponding information is unavalable. In particular, the works of regional VO's will be blocked. Is it possible to install LFC on separate computers and how?


    5. (ROC SouthEast Europe): AEGIS reports that an easy fix for the longstanding problems with the superficial gCE SAM failures is now available at: http://listserv.cclrc.ac.uk/cgi-bin/webadmin?A1=ind0705&L=lcg-rollout#6 .
      We will add that on a ticket / savannnah bug


    6. (ROC SouthEast Europe): [INFORMATION] IL-BGU reported that some RPMs are missing from the SL4 Compatibility repository: GGUS ticket 22183 (https://gus.fzk.de/ws/ticket_info.php?ticket=22183). We have also updated our wiki page with extra info about SL4 on production more info at http://wiki.egee-see.org/index.php/SL4_64bit_WN.


    7. (ROC SouthEast Europe): [INFORMATION] In SEE ROC we are starting to move small sites to use SL4 WNs in order to test regional apps.


  • 16:30 17:00
    WLCG Items 30m
    Speaker: Mr Daniele Bonacorsi (CNAF-INFN BOLOGNA, ITALY)
  • <big> LHCb service </big>
    1. LHCb have lost quite a lot of time (~a week) because the LFC instances at CERN have been upgraded without any prior announcement. In more details: last week the LFC has been upgraded to a newest version that doesn't allow queries in case of VOMS extensions expired. LHC-b kept however running reconstruction jobs using proxy (long living ones) whose voms extension were expired.
      By the way, voms-proxy-init command should present a consistent behavior and shouldn't allow to have different lifetimes for flat and voms extension component.
    2. In the aim of gettig a better SRM servicefrom our T1, LHCb have integrated specific SRM checks on their SAM suite.
      We would like to inform all site managers the LHCb will engage in a campain of close monitoring of our endpoints by using the results of these specific tests.
      We also would like to inform that these tests are set as critical and will affect the LHCb specific site availaility calculation.
      A preliminary set of tests are already available at: http://santinel.home.cern.ch/santinel/cgi-bin/srm_test further tests likedirect file access are coming soon.
    Speaker: Dr roberto santinelli (CERN/IT/GD)
  • <big> ALICE service </big>
    Speaker: Dr Patricia Mendez Lorenzo (CERN IT/GD)
  • <big> WLCG Service Coordination Issues </big>
    WLCG Collaboration workshop September 1-2 2007, Victoria, BC, Canada (co-located with CHEP 2007)
    Speaker: Jamie Shiers / Harry Renshall
  • 16:55 17:00
    OSG Items 5m
    1. Item 1
  • 17:00 17:05
    Review of action items 5m
    list of actions
  • 17:10 17:15
    AOB 5m
  • Operations workshop in Stockholm, 13-15th June, agenda available:
    http://indico.cern.ch/conferenceTimeTable.py?confId=12807