ATLAS UK Cloud Support

Europe/London
Zoom

Zoom

Tim Adye (Science and Technology Facilities Council STFC (GB)), James William Walder (Science and Technology Facilities Council STFC (GB))
Description

https://cern.zoom.us/j/98434450232

Password protected (same as (new) OPs Mtg)

Outstanding tickets

  • 150277 UKI-LT2-QMUL less urgent in progress 2021-01-20 16:35:00 UKI-LT2-QMUL: Transfer issues and high efficiency error in the past 12 hrs
    • Transfer failures; to follow up
  • 150252 UKI-NORTHGRID-MAN-HEP less urgent in progress 2021-01-19 06:52:00 UKI-NORTHGRID-MAN-HEP has efficiency errors
    • In progress, may be related to (time-localised) uk-wide drop of transfer eff?
  • 149842 UKI-SCOTGRID-ECDF very urgent in progress 2021-01-19 10:43:00 UKI-SCOTGRID-ECDF: Low transfer efficiency due to TRANSFER ERROR: Copy failed with mode 3rd pull, wi…
    • In progress
  • 149362 UKI-SOUTHGRID-RALPP urgent in progress 2021-01-05 12:35:00 ATLAS CE failures on UKI-SOUTHGRID-RALPP-heplnx207
    • Stalled - awaiting input from Atlas experts
  • 148342 UKI-SCOTGRID-GLASGOW less urgent in progress 2021-01-17 21:34:00 UKI-SCOTGRID-GLASGOW with transfer efficiency degraded and many failures
    • DPM - to be investigated
    • Ceph issues; Using Xrootd spaces (ie. rather than giving it a big space):
      • Issues in internal cache
      • Unable to write new symlinks
      • Due to how the ‘filesystem’ deals with rucio ‘hashed’ paths.
        • Error message same as different error source; difficult to debug. Should not re-occur, now that links for all paths created.
  • 146651 RAL-LCG2 urgent on hold 2021-01-19 10:05:00 singularity and user NS setup at RAL
    • No progress; other VOs starting to make requests.
  • 142329 UKI-SOUTHGRID-SUSX top priority on hold 2021-01-20 20:29:00 CentOS7 migration UKI-SOUTHGRID-SUSX
    • Ticket updated.

CPU

  • RAL

    • CE01 failure; Only atlas jobs dropped; now recovering and reclaiming slots from CMS.
    • Same problem as over Christmas; but different CE
    • Is not expected to stop submitting on all CE’s if one CE stops
  • Northgrid

    • Lancs - familiar full disk servers / load-balancing hand-holding needed.
  • London

    • QMUL corrected Storm space usage, and added quota (done today).
      • Howver still above watermark (as ATLAS wrote more data);
      • stopping jobs from completing (no space to write)
      • Rucio now started deletions; now that space is correct.
  • SouthGrid

    • Ox offline yesterday for arc-CE downgrade; complete.
  • Scotgrid

    • Durham; old workernodes might have bad disks
      • Update - looks like HC test file is missing; to declare lost, but need to follow-up
      • not the first time it’s observed at Durham.

Other new issues

Ongoing issues

  • CentOS7 - Sussex

    • As ticket
  • TPC with http

    • NTR
  • Storageless Site test / storage decomissioning (Oxford)

    • Progress on CE, needs some input from Sam for next steps
  • ECDF volatile storage

    • JW started to look; many steps need Rucio / DMM core experts to implement
    • JW to bump this forward
  • Glasgow DPM Decommissioning

    • Activity still ongoing for final steps
  • ATLAS: Site Availability/Reliability reports: Glasgow

    • No news on CRIC migration; JW to follow-up

News round-table

  • Vip
  • Downtime for arc downgrade
    • For the Xcache will follow-up with Sam for help
  • Dan
    • Storm; db stops being updated; for used space.
    • Now corrected, but ATLAS has been filling up and now over quota
      • JW: Rucio now noticed and begun deletions
  • Matt
    • Still hand-holding activities with DPM.
  • Peter
    • NTR
  • Sam
    • ntr
  • Gareth
    • 4-500 jobs; unspecified gridmanager error; analysis errors due to excess memory?
    • Discussion on numbers of gridFTP connections that are acceptable from Site;
      • Agreed that setting a reasonable max number of gridFTP connections is sensible
      • JW to provide RAL config settings for this
  • JW
    • ADC meeting to have new Cloud section; hope for better possibilities to raise non-urgent issues upwards
  • Duncan
    • NTR
  • Patrick
    • NTR

AOB

There are minutes attached to this event. Show them.
    • 10:00 10:20
      Status 20m
      • 150277 UKI-LT2-QMUL less urgent in progress 2021-01-20 16:35:00 UKI-LT2-QMUL: Transfer issues and high efficiency error in the past 12 hrs
        • Transfer failures; to follow up
      • 150252 UKI-NORTHGRID-MAN-HEP less urgent in progress 2021-01-19 06:52:00 UKI-NORTHGRID-MAN-HEP has efficiency errors
        • In progress, may be related to (time-localised) uk-wide drop of transfer eff?
      • 149842 UKI-SCOTGRID-ECDF very urgent in progress 2021-01-19 10:43:00 UKI-SCOTGRID-ECDF: Low transfer efficiency due to TRANSFER ERROR: Copy failed with mode 3rd pull, wi…
        • In progress
      • 149362 UKI-SOUTHGRID-RALPP urgent in progress 2021-01-05 12:35:00 ATLAS CE failures on UKI-SOUTHGRID-RALPP-heplnx207
        • Stalled - awaiting input from Atlas experts
      • 148342 UKI-SCOTGRID-GLASGOW less urgent in progress 2021-01-17 21:34:00 UKI-SCOTGRID-GLASGOW with transfer efficiency degraded and many failures
        • DPM - to be investigated
        • Ceph issues; Using Xrootd spaces (ie. rather than giving it a big space):
          • Issues in internal cache
          • Unable to write new symlinks
          • Due to how the ‘filesystem’ deals with rucio ‘hashed’ paths.
            • Error message same as different error source; difficult to debug. Should not re-occur, now that links for all paths created.
      • 146651 RAL-LCG2 urgent on hold 2021-01-19 10:05:00 singularity and user NS setup at RAL
        • No progress; other VOs starting to make requests.
      • 142329 UKI-SOUTHGRID-SUSX top priority on hold 2021-01-20 20:29:00 CentOS7 migration UKI-SOUTHGRID-SUSX
        • Ticket updated.
      • Outstanding tickets 10m
      • CPU 5m

        New link for the site-oriented dashboard

        • RAL

          • CE01 failure; Only atlas jobs dropped; now recovering and reclaiming slots from CMS.
          • Same problem as over Christmas; but different CE
          • Is not expected to stop submitting on all CE’s if one CE stops
        • Northgrid

          • Lancs - familiar full disk servers / load-balancing hand-holding needed.
        • London

          • QMUL corrected Storm space usage, and added quota (done today).
            • Howver still above watermark (as ATLAS wrote more data);
            • stopping jobs from completing (no space to write)
            • Rucio now started deletions; now that space is correct.
        • SouthGrid

          • Ox offline yesterday for arc-CE downgrade; complete.
        • Scotgrid

          • Durham; old workernodes might have bad disks
            • Update - looks like HC test file is missing; to declare lost, but need to follow-up
            • not the first time it’s observed at Durham.
      • Other new issues / tasks 5m
    • 10:20 10:40
      Ongoing Items 20m
      • CentOS7 - Sussex

        • As ticket
      • TPC with http

        • NTR
      • Storageless Site test / storage decomissioning (Oxford)

        • Progress on CE, needs some input from Sam for next steps
      • ECDF volatile storage

        • JW started to look; many steps need Rucio / DMM core experts to implement
        • JW to bump this forward
      • Glasgow DPM Decommissioning

        • Activity still ongoing for final steps
      • ATLAS: Site Availability/Reliability reports: Glasgow

        • No news on CRIC migration; JW to follow-up
    • 10:40 10:50
      News round-table 10m
      • Vip
        • Downtime for arc downgrade
          • For the Xcache will follow-up with Sam for help
      • Dan
        • Storm; db stops being updated; for used space.
        • Now corrected, but ATLAS has been filling up and now over quota
          • JW: Rucio now noticed and begun deletions
      • Matt
        • Still hand-holding activities with DPM.
      • Peter
        • NTR
      • Sam
        • ntr
      • Gareth
        • 4-500 jobs; unspecified gridmanager error; analysis errors due to excess memory?
        • Discussion on numbers of gridFTP connections that are acceptable from Site;
          • Agreed that setting a reasonable max number of gridFTP connections is sensible
          • JW to provide RAL config settings for this
      • JW
        • ADC meeting to have new Cloud section; hope for better possibilities to raise non-urgent issues upwards
      • Duncan
        • NTR
      • Patrick
        • NTR
    • 10:50 11:00
      AOB 10m