US ATLAS Computing Integration and Operations
virtual room
your office
-
-
13:00
→
13:15
Top of the Meeting 15mSpeakers: Michael Ernst, Robert William Gardner Jr (University of Chicago (US))
-
Caching and the Cloudy Tier2 Program 5mSpeaker: Robert William Gardner Jr (University of Chicago (US))
-
-
13:15
→
13:25
Production 10mSpeakers: Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US))
-
13:25
→
13:30
Data Management 5mSpeaker: Armen Vartapetian (University of Texas at Arlington (US))
-
13:30
→
13:35
Data transfers 5mSpeaker: Hironori Ito (Brookhaven National Laboratory (US))
-
13:35
→
13:40
Networks 5mSpeaker: Dr Shawn McKee (University of Michigan ATLAS Group)
-
13:40
→
13:45
FAX 5mSpeakers: Ilija Vukotic (University of Chicago (US)), Wei Yang (SLAC National Accelerator Laboratory (US))
-
13:45
→
13:55
Updates on Leadership HPC Integration and Event Service 10mSpeaker: Torre Wenaus (Brookhaven National Laboratory (US))
-
14:05
→
15:05
Site Reports
-
14:05
BNL 5mSpeaker: Michael Ernst
- Experienced dCache storage server performance problems in the last few days. Investigation unveiled that those were caused by production jobs reading data via dCap, which was (unexpectedly) proxied through NFS4.1. The latter caused high load on the storage servers, impacting the read and write performance. Solved by configuring dCap access via URL-style addressing in the site mover.
- FTS3 issue. Receiving a large number (~300k) of tape staging requests from Rucio the processing time per request within FTS increased to >20 minutes, causing the majority of them to time out. It was found that some staging-related database queries were the reason. The FTS database was reshuffled (dump & restore) and indices were configured to speed up the slow queries. This has helped to improve the overall FTS performance.
-
14:10
AGLT2 5mSpeakers: Robert Ball (University of Michigan (US)), Dr Shawn McKee (University of Michigan ATLAS Group)
Thirteen R630 have been received at UM, and 10 at MSU. The UM servers are under test before being placed into production. Work on PDUs is in progress at MSU before the servers can be installed.
We have determined that it is critical to the performance of the R630 and the R730 (as much as 15%) to correctly configure the memory. DIMMs should be installed in multiples of 4, with each processor bank identically populated, else it does not access at the highest efficiency. This is because there are 4 memory channels that must be equally populated. As a result of this discovery we are re-distributing memory, and will have some systems with 128GB, and some with 256GB, instead of having all system with 192GB. With Hyper-Threading, these have 48 cores that will be Dynamically configured in Condor.
An overload of the /var partition on our gate04 gatekeeper on Friday caused a loss of all running jobs. The source of this is understood (spooling of files) and is under control.
The proddisk token no longer has any assigned space, and all associated dark data was deleted. Space was moved to the datadisk token.
-
14:15
MWT2 5mSpeakers: David Lesny (Univ. Illinois at Urbana-Champaign (US)), Lincoln Bryant (University of Chicago (US))
- New hardware
- UC - Servers racked. New line card installed into switch.
- UIUC - Servers have been ordered
- IU - Order should be placed in a few days
- NSS/NSPR update pushed to all nodes
- HTCondorCE fix for Collector runaway memory usage
- Collector memory usage grows infinitely killing the CE
- Add this to /etc/condor-ce/config/99-local.conf
- GSS_ASSIST_GRIDMAP_CACHE_EXPIRATION = 1800
- CA deployment cleanup on SE for new CAs
- CVMFS 2.1.20
- Server Nearly deployed
- Client fully deployed
- PortableCMFS updated
- PRODDISK Decommission nearly complete
- All datasets removed
- Removing all darkdata
- UIUC in downtime
- ICC PM
- Relocating all hardware for new "pod" deployments
- New hardware
-
14:20
NET2 5mSpeaker: Prof. Saul Youssef (Boston University (US))
Smooth running.
We were unaware of the NSS update, but will follow up asap.
PRODDISK migration essentially completed globally about a couple of weeks ago http://egg.bu.edu//ATLAS%7Binf:ATLAS%7D/gadget:Schedconfig/section:report/%7Bsubsection:PRODDISK%20migration,cron:2%20hour%7D/index.html
We are collecting quotes and preparing for this year's hardware purchases. This will include at least 1.1 PB of storage and worker nodes. I'll be in touch.
There are many things happening in parallel, but a couple of issues might be interesting for other sites:
o Two ADC meetings ago, instructions were given to set up sites for central dark data detection. We're preparing that.
o Because central deletion doesn't delete directories, we are slowly accumulating empty rucio directories which are likely appearing at all the sites. There are at most 65,000 such per scope, but since there is one scope per user, this can eventually reach 65 M empty directories. It's very likely that we can just delete old empty rucio directories, but we're going to confirm with ddm first.
- 14:25
-
14:30
SWT2-UTA 5mSpeaker: Patrick Mcguigan (University of Texas at Arlington (US))
UTA_SWT2
- Updated nss and nspr rpms, but a cluster command caused an issue on certain nodes that broke the gridftp server and caused some errors (since resolved).
- Will take a downtime on Friday 11/27 thru 11/30 as the facility is undergoing major electrical upgrades
SWT2_CPB
- Updated nss and nspr rpms, no problems discovered at this point
- We arranging some electrical work in the machine room to add more power drops to fill out our current breaker panels.
- Starting to work on a purchase of compute nodes to close out FY15 funds. Expect on the order of 30 compute nodes
-
14:35
WT2 5mSpeaker: Wei Yang (SLAC National Accelerator Laboratory (US))
* 11 blade servers arrived. each 24 physical cores/128GB/1TB. Top of the rack switch hasn't arrive. Will test OpenStack first. Expect to be in production in Jan-Feb.
* Ordering 2x 84bay storage from RAID Inc. with 8TB SMR drives. Still have to negotiate for 5yr support. RAID Inc. currently provides 3yr support.
* Work with OSG to sort out the WLCG accounting report issue. Looks like it is mostly resolved.
* CondorCE problem still going on. Lost ~5% of jobs everyday (CondorCE marks the jobs as completed, clean the disk space (including the x509 proxies needed by the job), but leave the LSF jobs running). The Condor team promised to work on it.
-
14:05
-
15:05
→
15:10
AOB 5m
-
13:00
→
13:15