Outstanding tickets
- 150277 UKI-LT2-QMUL less urgent in progress 2021-01-20 16:35:00 UKI-LT2-QMUL: Transfer issues and high efficiency error in the past 12 hrs
- Transfer failures; to follow up
- 150252 UKI-NORTHGRID-MAN-HEP less urgent in progress 2021-01-19 06:52:00 UKI-NORTHGRID-MAN-HEP has efficiency errors
- In progress, may be related to (time-localised) uk-wide drop of transfer eff?
- 149842 UKI-SCOTGRID-ECDF very urgent in progress 2021-01-19 10:43:00 UKI-SCOTGRID-ECDF: Low transfer efficiency due to TRANSFER ERROR: Copy failed with mode 3rd pull, wi…
- 149362 UKI-SOUTHGRID-RALPP urgent in progress 2021-01-05 12:35:00 ATLAS CE failures on UKI-SOUTHGRID-RALPP-heplnx207
- Stalled - awaiting input from Atlas experts
- 148342 UKI-SCOTGRID-GLASGOW less urgent in progress 2021-01-17 21:34:00 UKI-SCOTGRID-GLASGOW with transfer efficiency degraded and many failures
- DPM - to be investigated
- Ceph issues; Using Xrootd spaces (ie. rather than giving it a big space):
- Issues in internal cache
- Unable to write new symlinks
- Due to how the ‘filesystem’ deals with rucio ‘hashed’ paths.
- Error message same as different error source; difficult to debug. Should not re-occur, now that links for all paths created.
- 146651 RAL-LCG2 urgent on hold 2021-01-19 10:05:00 singularity and user NS setup at RAL
- No progress; other VOs starting to make requests.
- 142329 UKI-SOUTHGRID-SUSX top priority on hold 2021-01-20 20:29:00 CentOS7 migration UKI-SOUTHGRID-SUSX
CPU
-
RAL
- CE01 failure; Only atlas jobs dropped; now recovering and reclaiming slots from CMS.
- Same problem as over Christmas; but different CE
- Is not expected to stop submitting on all CE’s if one CE stops
-
Northgrid
- Lancs - familiar full disk servers / load-balancing hand-holding needed.
-
London
- QMUL corrected Storm space usage, and added quota (done today).
- Howver still above watermark (as ATLAS wrote more data);
- stopping jobs from completing (no space to write)
- Rucio now started deletions; now that space is correct.
-
SouthGrid
- Ox offline yesterday for arc-CE downgrade; complete.
-
Scotgrid
- Durham; old workernodes might have bad disks
- Update - looks like HC test file is missing; to declare lost, but need to follow-up
- not the first time it’s observed at Durham.
Other new issues
Ongoing issues
-
CentOS7 - Sussex
-
TPC with http
-
Storageless Site test / storage decomissioning (Oxford)
- Progress on CE, needs some input from Sam for next steps
-
ECDF volatile storage
- JW started to look; many steps need Rucio / DMM core experts to implement
- JW to bump this forward
-
Glasgow DPM Decommissioning
- Activity still ongoing for final steps
-
ATLAS: Site Availability/Reliability reports: Glasgow
- No news on CRIC migration; JW to follow-up
News round-table
- Vip
- Downtime for arc downgrade
- For the Xcache will follow-up with Sam for help
- Dan
- Storm; db stops being updated; for used space.
- Now corrected, but ATLAS has been filling up and now over quota
- JW: Rucio now noticed and begun deletions
- Matt
- Still hand-holding activities with DPM.
- Peter
- Sam
- Gareth
- 4-500 jobs; unspecified gridmanager error; analysis errors due to excess memory?
- Discussion on numbers of gridFTP connections that are acceptable from Site;
- Agreed that setting a reasonable max number of gridFTP connections is sensible
- JW to provide RAL config settings for this
- JW
- ADC meeting to have new Cloud section; hope for better possibilities to raise non-urgent issues upwards
- Duncan
- Patrick
AOB
There are minutes attached to this event.
Show them.