US ATLAS Computing Integration and Operations
-
-
13:00
→
13:05
Top of the Meeting 5mSpeakers: Eric Christian Lancon (CEA/IRFU,Centre d'etude de Saclay Gif-sur-Yvette (FR)), Robert William Gardner Jr (University of Chicago (US))
-
13:10
→
13:20
Capacity News: Procurements & Retirements 10m
-
13:25
→
13:26
Sites acknowledgement at the ATLAS week 1mSpeaker: Robert Ball (University of Michigan (US))
There was a central request from management (I have not found the original request) to make sure that support is correctly attributed to the funding agencies involved and that they can easily be recognized. This is also tied to putting the correct NAME to a site.
The AGLT2 name, in OIM, is relatively simple. BNL is very complicated. There are multiple CEs active, and varying names, or none(?) seem to be in use. I had suggested using the "Accounting Name" field that can be stripped from public-available XML. But then it was just pushed and new fields were placed in AGIS. It is done, but everyone needs to check was was put in place. The sites "Recognition Name", known below as "Alternative Name", was placed in this way, but the results may not be satisfactory for your site.
Each "Site" (start at http://atlas-agis.cern.ch/agis/site/table_view/ and select your site from column 1, or start at http://atlas-agis.cern.ch/agis/atlassite/table_view/?vo_name=atlas and select your site name from the 5th column) now has an "Alternative name" field at the top. Please check yours THIS WEEK and make sure it is what you want; if not, then correct it, or send to me to ask to correct or, or just send to me what you want and we will make sure it is set that way.
There is also a TimeZone field that should be filled.
Note that the pre-assigned names for US T2s in particular may be a bit funky, especially when there are multiple, distinct sites as is the case with the SWT2.
NOTE: If you make a change, or if you are happy, either way, LET ME KNOW. I have a spread-sheet to maintain.
-
13:35
→
13:45
Production 10mSpeaker: Mark Sosebee (University of Texas at Arlington (US))
1) MC15c almost done (>4B events)
2) MC16_13TeV in preparation
3) Most sites generally full - some not quite (due to 1)?)
4) Work on transitioning to RH7 - see:
https://twiki.cern.ch/twiki/bin/view/AtlasComputing/CentOS7Readiness
https://its.cern.ch/jira/browse/ADCINFR-11
Note: ATLAS will not mass-migrate until LS2
-
13:45
→
13:50
Data Management 5mSpeaker: Armen Vartapetian (University of Texas at Arlington (US))
-
13:50
→
13:55
Data transfers 5mSpeaker: Hironori Ito (Brookhaven National Laboratory (US))
-
13:55
→
14:00
Networks 5mSpeaker: Dr Shawn McKee (University of Michigan ATLAS Group)
-
14:00
→
14:05
FAX and Xrootd Caching 5mSpeakers: Andrew Bohdan Hanushevsky (SLAC National Accelerator Laboratory (US)), Andrew Hanushevsky (STANFORD LINEAR ACCELERATOR CENTER), Andrew Hanushevsky, Ilija Vukotic (University of Chicago (US)), Wei Yang (SLAC National Accelerator Laboratory (US))
-
14:25
→
15:25
Site Reports
-
14:25
BNL 5mSpeaker: Eric Christian Lancon (CEA/IRFU,Centre d'etude de Saclay Gif-sur-Yvette (FR))
-
14:30
AGLT2 5mSpeakers: Robert Ball (University of Michigan (US)), Dr Shawn McKee (University of Michigan ATLAS Group)
In advance of storage upgrades we have retired one MD1200 shelf of 2TB disks for spares as those had become depleted.
Working with HTCondor people (Greg Thain) about the periodic misses we see in our job counting on the gatekeepers using condor_q. There is some similarity to an issue found at BNL, and that will be corrected in release 8.4.7. Greg has made available pre-release copies of the 8.4.7 rpms that we are currently trying out (without apparent problems) on our test gatekeeper, gate03.
-
14:35
MWT2 5mSpeakers: David Lesny (Univ. Illinois at Urbana-Champaign (US)), Lincoln Bryant (University of Chicago (US))
Site has been running well
Illinois Campus Cluster down for monthly PM
- Main goal is upgrading GPFS to V4
- Future downtime to migration file system to new format
Testing CVMFS 2.2.1
- 2.2.2 has been release for use however...
- Bug fix for cache issue in 2.2.3
Increase memory on Perfsonar nodes to 16GB
Ceph setup for logs and event service
- Entries in AGIS are configured
- Some small level testing has been done
- Lincoln is investigating a large scale test
Docker for Analysis preservation
- Request from NYU for Docker service to scale up analysis preservation
- Containerize a number of MWT2 job slots suitable for RECAST
Upgrades @ Illinois
- Getting prices from the ICC with Broadwell based nodes
- Need to move all equipment to new racks in June
- 10Gb top of rack switches along with 1Gb IPIMI switches
Upgrade @ UChicago
- Getting prices for additional hypervisors
- Getting prices for additional dCache servers in order to add space token capacity
-
14:40
NET2 5mSpeaker: Prof. Saul Youssef (Boston University (US))
- 14:45
-
14:50
SWT2-UTA 5mSpeaker: Patrick Mcguigan (University of Texas at Arlington (US))
1) Beginning discussions with Dell regarding possible storage and compute node purchases in response to potential for additional funding. Phone call with Dell on Friday, 5/27.
2) Issue with the cooling in the machine room for SWT2_CPB has hopefully been resolved. Facilities Management / A/C shop informed us that changes were implemented to stabilize the flow of chilled water to the CRAC units. Temperature data over the past week look very good.
3) Expecting to begin testing the LHCONE/DMZ network link the week of 5/30.
- 14:55
-
14:25
-
15:25
→
15:30
AOB 5m
-
13:00
→
13:05