ATLAS UK Cloud Support

Europe/London
Vidyo

Vidyo

Tim Adye (Science and Technology Facilities Council STFC (GB)), James William Walder (Science and Technology Facilities Council STFC (GB))

Outstanding tickets

  • 147979 UKI-NORTHGRID-MAN-HEP less urgent assigned 2020-07-27 07:48:00 UKI-NORTHGRID-MAN-HEP timeout transfer errros and also deletion errors

    • Related to lost disk server; (there no ‘maintaince mode’, so deletions are attempted while data is being recovered); firewalling helped to mitigate this.
  • 146771 UKI-SCOTGRID-ECDF less urgent in progress 2020-07-29 12:15:00 UKI-SCOTGRID-ECDF deletion failures with “The requested service is not available at the moment.”

    • long update provide in the GGUS; restarting xroot frequently on head node.
    • Uncertain why this seems particular to ECDF:
      • Issues appears related to Xrootd (and server), e.g tcp reuse
      • following up with developers
  • 146651 RAL-LCG2 urgent on hold 2020-07-23 20:02:00 singularity and user NS setup at RAL

    • on hold
  • 146374 UKI-NORTHGRID-SHEF-HEP urgent in progress 2020-07-22 14:53:00 ATLAS pilot jobs idle on UKI-NORTHGRID-SHEF-HEP CE

    • no news
  • 144759 UKI-SCOTGRID-GLASGOW less urgent on hold 2020-06-09 07:59:00 High traffic from UKI-SCOTGRID-GLASGOW on RAL CVMFS Stratum1

    • on hold
  • 142329 UKI-SOUTHGRID-SUSX top priority on hold 2020-06-04 14:05:00 CentOS7 migration UKI-SOUTHGRID-SUSX

    • on hold

CPU

  • RAL

    • no issues
  • Northgrid

    • Small dips for MAN (disk related);
    • LANCS remains below full capacity, with arc-ce problems getting pilots to run; and disk server issues (likely ipv6 related)
  • London

    • Power glitch at QMUL ; (Sites set to DT afterwards)
    • Aim to resolve all residual issues today
  • SouthGrid

    • CAM issues (related to QMUL)
  • Scotgrid

    • brief dip for Durham; may just be from other VO demands

Other new issues

Ongoing issues

  • CentOS7 - Sussex

    • On Hold
  • Grand Unified queues

    • On Hold

News round-table

  • Vip

    • NTR
  • Dan

    • Aim to fix residual problems from power issue today
  • Matt

    • NTR
  • Gareth

    • Will continue to slowly add cores (in 128 job slots (64 core) blocks)
  • JW

    • NTR

 

 

There are minutes attached to this event. Show them.
    • 10:00 10:20
      Status 20m
      • Outstanding tickets 10m
        • 147979 UKI-NORTHGRID-MAN-HEP less urgent assigned 2020-07-27 07:48:00 UKI-NORTHGRID-MAN-HEP timeout transfer errros and also deletion errors

          • Related to lost disk server; (there no ‘maintaince mode’, so deletions are attempted while data is being recovered); firewalling helped to mitigate this.
        • 146771 UKI-SCOTGRID-ECDF less urgent in progress 2020-07-29 12:15:00 UKI-SCOTGRID-ECDF deletion failures with “The requested service is not available at the moment.”

          • long update provide in the GGUS; restarting xroot frequently on head node.
          • Uncertain why this seems particular to ECDF:
            • Issues appears related to Xrootd (and server), e.g tcp reuse
            • following up with developers
        • 146651 RAL-LCG2 urgent on hold 2020-07-23 20:02:00 singularity and user NS setup at RAL

          • on hold
        • 146374 UKI-NORTHGRID-SHEF-HEP urgent in progress 2020-07-22 14:53:00 ATLAS pilot jobs idle on UKI-NORTHGRID-SHEF-HEP CE

          • no news
        • 144759 UKI-SCOTGRID-GLASGOW less urgent on hold 2020-06-09 07:59:00 High traffic from UKI-SCOTGRID-GLASGOW on RAL CVMFS Stratum1

          • on hold
        • 142329 UKI-SOUTHGRID-SUSX top priority on hold 2020-06-04 14:05:00 CentOS7 migration UKI-SOUTHGRID-SUSX

          • on hold
      • CPU 5m
        • RAL

          • no issues
        • Northgrid

          • Small dips for MAN (disk related);
          • LANCS remains below full capacity, with arc-ce problems getting pilots to run; and disk server issues (likely ipv6 related)
        • London

          • Power glitch at QMUL ; (Sites set to DT afterwards)
          • Aim to resolve all residual issues today
        • SouthGrid

          • CAM issues (related to QMUL)
        • Scotgrid

          • brief dip for Durham; may just be from other VO demands

         

         

      • Other new issues 5m
    • 10:20 10:40
      Ongoing issues 20m
      • CentOS7 - Sussex

        • On Hold
      • Grand Unified queues

        • On Hold
    • 10:40 10:50
      News round-table 10m
      • Vip

        • NTR
      • Dan

        • Aim to fix residual problems from power issue today
      • Matt

        • NTR
      • Gareth

        • Will continue to slowly add cores (in 128 job slots (64 core) blocks)
      • JW

        • NTR
    • 10:50 11:00
      AOB 10m