US ATLAS Computing Facility

US/Eastern
    • 13:00 13:10
      WBS 2.3 Facility Management News 10m
      Speakers: Eric Christian Lancon (BNL), Robert William Gardner Jr (University of Chicago (US))

      - Setup a test instance somewhere on US facility for testing OSG new releases

      - Presentations at forthcoming meetings

      • BOINC backfilling (MSU)
      • Ceph deployment, operation, performances (NET2)

       

       

    • 13:10 13:20
      OSG-LHC 10m
      Speakers: Brian Lin (University of Wisconsin), Matyas Selmeci

      3.4.27 (tentatively next week)

      Looking for testers for the following (all available in osg-testing):

      • XRootD 4.9.1 RC2 (https://github.com/xrootd/xrootd/issues/937
      • globus-ftp-client-9.1-2.1 (changed sources from GT to GCT)
      • globus-gridftp-server-13.9-1.1 (changed sources from GT to GCT)
      • myproxy-6.2.3-1.1 (changed sources from GT to GCT)
      • HTCondor-CE 3.2.2 (https://github.com/opensciencegrid/htcondor-ce/releases/tag/v3.2.2)
      • CVMFS 2.6.0 (https://cvmfs.readthedocs.io/en/2.6/cpt-releasenotes.html)
    • 13:20 13:40
      Topical Report
    • 13:40 14:25
      US Cloud Status
      • 13:40
        US Cloud Operations Summary 5m
        Speaker: Mark Sosebee (University of Texas at Arlington (US))
      • 13:45
        BNL 5m
        Speaker: Xin Zhao (Brookhaven National Laboratory (US))
        • migration to SL7
          • test ongoing
          • tentative target is April (rolling fashion)
        • setup ANALY PQ to use GPUs on BNL IC
          • ongoing, timeline 3~4 weeks
      • 13:50
        AGLT2 5m
        Speakers: Philippe Laurens (Michigan State University (US)), Dr Shawn McKee (University of Michigan ATLAS Group), Prof. Wenjing Wu (Computer Center, IHEP, CAS)

        1. Ucore has been running for a week smoothly and uses about 30% of the cores, with a mix of 1 core and 8 core jobs.
        2. For about a week, there has been a shortage of 1 core job, so the cluster wall time utilization dropped to about 93%, otherwise it was 99.5%. This is because we still have 17% static slots configured for 1 core job. http://gate01.aglt2.org/condorq.html

        3. We had a glitch with dCache xrootd doors, primarily because of the expiration/invalidation of a proxy certificate for the dcache user account, Now fixed..

        4. We discovered that BOINC jobs were occasionally causing some problems in the /tmp area. Ghost processes (jobs killed by the sanity check scripts) were not releasing big deleted files.
        We had 2 tickets about jobs failing with not having enough space. We know how to manual recover and are working on a long term solution..
        5. We are continuing to improve our custom scripts checking the "health" of the woker nodes to automatically start and retire condor.
        6. MSU site had an air conditioner failure from a shorted fan motor. That meant one emergency power off of  a fraction of the worker nodes last week for about an hour until that fan was bypassed.  The fan was just replaced. Back to normal.

         

      • 13:55
        MWT2 5m
        Speakers: Judith Lorraine Stephen (University of Chicago (US)), Lincoln Bryant (University of Chicago (US))

        Brief UC/IU worker downtime last week, resolved now

        UC

        • Working on getting the new GPU machine online
        • Looking into the reported file corruption issues reported yesterday

        IU

        • Fred and Neeha are in the process of getting the twelve new C6420s online
        • Brief site issue on the workers after IPv6 was enabled on the IU switches
          • Workers were trying to access CERN via v6
          • Disabled IPv6 on the IU workers for now

        UIUC

        • Compute order placed, should ship this week
      • 14:00
        NET2 5m
        Speaker: Prof. Saul Youssef (Boston University (US))

         

        Updating OSG so we can retire LSM and use rucio mover, then RH7 upgrade & singularity.

        NESE testing in progress.   Will setup ATLAS DDM endpoint soon. 

        Had to deal with GPFS hardware problem in system pool.  Need to evacuate and rebuild part of the pool. 

        New worker nodes online. 

      • 14:05
        SWT2 5m
        Speakers: Dr Horst Severini (University of Oklahoma (US)), Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US)), Patrick Mcguigan (University of Texas at Arlington (US))

        OU: Nothing to report, all running well.

               Not getting any UCORE/MCORE jobs, but apparently neither do most other US sites right now.

        UTA:

        • Discovered an issue with proxy lifetimes coming from BNL APF system.  HTCondor-CE negotiates a 24 hour proxy at submission time.  This can be overcome at the submit side with an extra directive.  Proxies are now OK.
        • UCORE testing is now started at UTA_SWT2.  We can duplicate the configuration at SWT2_CPB after testing.

         

      • 14:10
        HPC Operations 5m
        Speaker: Doug Benjamin (Duke University (US))

        Tadashi added new code to Harvester to get "left over" HITS files from shared file system and it works now at NERSC and ALCF.

        We can now run capacity level jobs (>= 802 nodes) at ALCF. This allows us to use overburn since we have already used up our allocation. 85M hours out of 80 M hours used.

        Now testing 1024 node jobs at NERSC. This will give us a 20% discount on the charged hours.

        OLCF has used more than 125% of allocation (104 Mhours used). Now at very low priority on ALCC. Investigating using Harvester but at lower priority.

      • 14:15
        Analysis Facilities - SLAC 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
      • 14:20
        Analysis Facilities - BNL 5m
        Speaker: William Strecker-Kellogg (Brookhaven National Lab)
    • 14:25 14:30
      AOB 5m