US ATLAS Computing Facility
-
-
13:00
→
13:10
WBS 2.3 Facility Management News 10mSpeakers: Eric Christian Lancon (BNL), Robert William Gardner Jr (University of Chicago (US))
-
13:10
→
13:20
OSG-LHC 10mSpeakers: Brian Lin (University of Wisconsin), Matyas Selmeci
3.4.27 (tentatively next week)
Looking for testers for the following (all available in osg-testing):
- XRootD 4.9.1 RC2 (https://github.com/xrootd/xrootd/issues/937
- globus-ftp-client-9.1-2.1 (changed sources from GT to GCT)
- globus-gridftp-server-13.9-1.1 (changed sources from GT to GCT)
- myproxy-6.2.3-1.1 (changed sources from GT to GCT)
- HTCondor-CE 3.2.2 (https://github.com/opensciencegrid/htcondor-ce/releases/tag/v3.2.2)
- CVMFS 2.6.0 (https://cvmfs.readthedocs.io/en/2.6/cpt-releasenotes.html)
-
13:20
→
13:40
Topical Report
-
13:20
Hammercloud update 5mSpeaker: Jose Caballero Bejar (Brookhaven National Laboratory (US))
-
13:20
-
13:40
→
14:25
US Cloud Status
-
13:40
US Cloud Operations Summary 5mSpeaker: Mark Sosebee (University of Texas at Arlington (US))
-
13:45
BNL 5mSpeaker: Xin Zhao (Brookhaven National Laboratory (US))
-
13:50
AGLT2 5mSpeakers: Philippe Laurens (Michigan State University (US)), Dr Shawn McKee (University of Michigan ATLAS Group), Prof. Wenjing Wu (Computer Center, IHEP, CAS)
1. Ucore has been running for a week smoothly and uses about 30% of the cores, with a mix of 1 core and 8 core jobs.
2. For about a week, there has been a shortage of 1 core job, so the cluster wall time utilization dropped to about 93%, otherwise it was 99.5%. This is because we still have 17% static slots configured for 1 core job. http://gate01.aglt2.org/condorq.html3. We had a glitch with dCache xrootd doors, primarily because of the expiration/invalidation of a proxy certificate for the dcache user account, Now fixed..
4. We discovered that BOINC jobs were occasionally causing some problems in the /tmp area. Ghost processes (jobs killed by the sanity check scripts) were not releasing big deleted files.
We had 2 tickets about jobs failing with not having enough space. We know how to manual recover and are working on a long term solution..
5. We are continuing to improve our custom scripts checking the "health" of the woker nodes to automatically start and retire condor.
6. MSU site had an air conditioner failure from a shorted fan motor. That meant one emergency power off of a fraction of the worker nodes last week for about an hour until that fan was bypassed. The fan was just replaced. Back to normal. -
13:55
MWT2 5mSpeakers: Judith Lorraine Stephen (University of Chicago (US)), Lincoln Bryant (University of Chicago (US))
Brief UC/IU worker downtime last week, resolved now
UC
- Working on getting the new GPU machine online
- Looking into the reported file corruption issues reported yesterday
IU
- Fred and Neeha are in the process of getting the twelve new C6420s online
- Brief site issue on the workers after IPv6 was enabled on the IU switches
- Workers were trying to access CERN via v6
- Disabled IPv6 on the IU workers for now
UIUC
- Compute order placed, should ship this week
-
14:00
NET2 5mSpeaker: Prof. Saul Youssef (Boston University (US))
-
14:05
SWT2 5mSpeakers: Dr Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))
OU: Nothing to report, all running well.
Not getting any UCORE/MCORE jobs, but apparently neither do most other US sites right now.
UTA:
- Discovered an issue with proxy lifetimes coming from BNL APF system. HTCondor-CE negotiates a 24 hour proxy at submission time. This can be overcome at the submit side with an extra directive. Proxies are now OK.
- UCORE testing is now started at UTA_SWT2. We can duplicate the configuration at SWT2_CPB after testing.
-
14:10
HPC Operations 5mSpeaker: Doug Benjamin (Duke University (US))
Tadashi added new code to Harvester to get "left over" HITS files from shared file system and it works now at NERSC and ALCF.
We can now run capacity level jobs (>= 802 nodes) at ALCF. This allows us to use overburn since we have already used up our allocation. 85M hours out of 80 M hours used.
Now testing 1024 node jobs at NERSC. This will give us a 20% discount on the charged hours.
OLCF has used more than 125% of allocation (104 Mhours used). Now at very low priority on ALCC. Investigating using Harvester but at lower priority.
-
14:15
Analysis Facilities - SLAC 5mSpeaker: Wei Yang (SLAC National Accelerator Laboratory (US))
-
14:20
Analysis Facilities - BNL 5mSpeaker: William Strecker-Kellogg (Brookhaven National Lab)
-
13:40
-
14:25
→
14:30
AOB 5m
-
13:00
→
13:10