RAL Tier1 Experiments Liaison Meeting

Europe/London
Access Grid (RAL R89)

Access Grid

RAL R89

Zoom Meeting ID
66811541532
Host
Alastair Dewhurst
Useful links
Join via phone
Zoom URL
    • 13:30 13:34
      Site Operations 4m
    • 13:34 13:35
      Experiment Operational Issues 1m
    • 13:35 13:40
      ATLAS Operations Report 5m
      Speakers: Brij Kishor Jashal (Rutherford Appleton Laboratory), Jyoti Prakash Biswal (Rutherford Appleton Laboratory)
    • 13:40 13:45
      CMS Operations Report 5m
      Speaker: Katy Ellis (Science and Technology Facilities Council STFC (GB))

      BAU.

      Jyothish recreated the VM-based AAA manager and RAL-redirector machines in proxmox. These went into production today. Some issue with the redirector failing SAM tests - we are fixing.

      Jyothish also testing the XCache on one of the recycled Antares machine which could eventually replace the proxy machines. 

      Issues seen with new accounting portal CPU efficiency calculation are fixed. Further checks required. 

      CMS are now using an 'overflow' method for CRAB jobs trying to target RAL T1 due to data being stored on Echo. These jobs are able to overflow to local (UK) T2s, but continue to read the data from Echo. This seems to be working fine, and I cannot see any obvious effect on the network. In theory this could have happened any time in the past on the whim of users...now it is being directed by CMS CRAB so will happen more frequently. Please let me know any issue with this.

      Still to deal with /unmerged/

      Sites are starting to ask what rates they should expect for DC27. I have met with ATLAS colleagues, and next week with WLCG Technical leads to discuss our high-level approach. Work has been done to re-assess the Run-4 rates with the latest information. Now we need to convert this to DC27 targets.

    • 13:45 13:50
      LHCb Operations Report 5m
      Speaker: Alexander Rogovskiy (Rutherford Appleton Laboratory)

      Operational issues:

      • Failed uploads from RAL WNs to ECHO last Friday
        • Potential overload from the merging +sprucing jobs (they are IO intensive) -- see plot attached
        • Checksum redirects to external gateways could have contributed to that.
      • DFC issues last week
        • Resulted in a few spikes of completed and rescheduled jobs
        • some mitigations applied, so hopefully it should work better now
      • High failure rate for Simulation jobs
        • Still present
        • The fix is ready, should be deployed with the next DIRAC release (hopefully next week).
    • 13:50 13:55
      ALICE Operations Report 5m
      Speaker: Alexander Rogovskiy (Rutherford Appleton Laboratory)

      NTR

    • 13:55 14:00
      LSST Operations Report 5m
      Speakers: Thomas Birkett, Timothy Noble (Science and Technology Facilities Council STFC (GB))
      • Cluster getting closer to being healthy again - some blacklisted modules from CVEs blocking services
      • LSST jobs running well:

        Some inconsistency between monitoring RAL monitoring - 36% efficiency
      • Near 100% from LSST montoring from HTCondor - difference in job efficiency vs Pilot efficiency?
      • Data movement to RAL 21TB this week
    • 14:00 14:01
      Tier-1 Projects 1m
    • 14:05 14:10
      XRootD Development 5m
      Speakers: Alexander Rogovskiy (Rutherford Appleton Laboratory), Jyothish Thomas (STFC)
    • 14:15 14:20
      Utilizing GPUs 5m
      Speakers: Dr Brij Kishor Jashal (Rutherford Appleton Laboratory), Thomas Birkett
    • 14:25 14:45
      Hardware deployment 20m

      Echo SN
      WN
      GPU

      Speaker: Jacob Ward
    • 14:45 14:46
      AOB 1m
    • 14:50 14:59
      Summary of Operational Status and Issues 9m
      Speakers: Brian Davies (Science and Technology Facilities Council STFC (GB)), Darren Moore, Thomas Birkett
    • 15:00 15:05
      Any other Business 5m
      Speakers: Brian Davies (Science and Technology Facilities Council STFC (GB)), Darren Moore