US ATLAS Computing Integration and Operations

US/Eastern
virtual room (your office)

virtual room

your office

Description
Notes and other material available in the US ATLAS Integration Program Twiki
    • 13:00 13:15
      Top of the Meeting 15m
      Speakers: Michael Ernst, Robert William Gardner Jr (University of Chicago (US))

      Michael - workshops up coming.  See email yesterday.

      • Renovating the WLCG with promising new technologies, esp. with eye to run 3 and run 4 (Ian Bird)
      • Forum - to work on these topics, to make informed decisions, working groups needed, etc.
      • Nominations to represent at the upcoming meeting (end Jan).  US Facilities role -- good opportunity to avoid isolating ourselves. 
      • WLCG Collaboration Workshop, Feb 1-3, 2016 in Lisbon
      • ATLAS Sites Jamboree Jan 27-29, 2016 at CERN

        • https://indico.cern.ch/event/440821/

          (the agenda is still evolving, typically many aspects relevant to facilities services incl. grid middleware, storage and compute are addressed)

       

       

      Accounting:

      • Questions from Dave regarding accounting
      • Wei, Patrick, Saul - have a look at local versus WLCG accounting

       

      Worker node benchmarking

       

      Broadwell processor - perhaps wait till April.

       

       

       

       

       

       

       

       

       

      • Revisiting Facility-wide deliverables from the US ATLAS Technical Planning Meeting 15m
        Speaker: Robert William Gardner Jr (University of Chicago (US))
    • 13:15 13:25
      Capacity News: Procurements & Retirements 10m

      See  http://bit.ly/usatlas-capacity  

       

       

       

    • 13:25 13:35
      Production 10m
      Speaker: Mark Sosebee (University of Texas at Arlington (US))
    • 13:35 13:40
      Data Management 5m
      Speaker: Armen Vartapetian (University of Texas at Arlington (US))
    • 13:40 13:45
      Data transfers 5m
      Speaker: Hironori Ito (Brookhaven National Laboratory (US))
    • 13:45 13:50
      Networks 5m
      Speaker: Dr Shawn McKee (University of Michigan ATLAS Group)
    • 13:50 13:55
      FAX and Xrootd Caching 5m
      Speakers: Ilija Vukotic (University of Chicago (US)), Wei Yang (SLAC National Accelerator Laboratory (US))
    • 14:15 15:15
      Site Reports
      • 14:15
        BNL 5m
        Speaker: Michael Ernst
        • At the T1 we will be retiring 3000 job slots provided by Nehalem based WNs procured in 2009. These nodes will stay in the data center but will be re-purposed as resources managed under OpenStack. 
        • We have started preparing the FY16 disk storage procurement
        • The storage management group is in the process of creating the inventory dumps requested by ADC. The process is delayed because of the ongoing deployment of the new storage hardware. 
        • We've noticed significantly higher than normal I/O (not CPU) load on our CVMFS proxies starting on December 2 (see attached plot)..From some debugging on the farm with systemtap, it appears that the issue may be caused by jobs in the BNL_PROD_MCORE queue.  It seems from Wednesday until Saturday morning (in the US) there were a number of jobs running in this queue which have spawned short-lived shell scripts and python code such as miniAthFile.py.  After some additional digging, it appears these may be Pile jobs, like [1].  Shortlived processes in the analysis queues are also generating accesses. In general, each time these shortlived processes are spwaned, we're seeing a significant amount of CVMFS access, likely due to LD_LIBRARY_PATH, and PATH settings, as well as the actual loading of the libraries/software involved in the execution.
      • 14:20
        AGLT2 5m
        Speakers: Robert Ball (University of Michigan (US)), Dr Shawn McKee (University of Michigan ATLAS Group)

        The requested storage consistency check scripts for all of our endpoints were completed on November 30, in time for central checks on the first of December.  The generated dumps were confirmed correct, and the AGLT2 ticket on this was subsequently closed.  These checks take place on head02, our dCache nameSpace server.

        All UM-purchased R630 (with one exception) are now up and running Condor jobs.  The last UM machine will be brought up later this week.  MSU-purchased machines are now being brought up, following successful reconfiguration of the rack and network switch architecture.  We expect these 10 machines to be online in Condor within the next week or so.  The v37 capacity spreadsheet has been updated to include the UM machines only at this time.

        A cross check of WLCG accounting for November, vs our own Condor accounting, showed agreement to within 1%, ie, WLCG showed 3,586,592hrs and Condor showed 3,554,081hrs.

        For the WLCG summary, see
        https://espace.cern.ch/WLCG-document-repository/Accounting/Tier-2/2015/november-15/Tier2_Accounting_Report_November2015.pdf

        Today saw the first big burst of real LMEM jobs at AGLT2.  At peak, we noted about 1100 jobs running, out of 1770 possible slots.

         

      • 14:25
        MWT2 5m
        Speakers: David Lesny (Univ. Illinois at Urbana-Champaign (US)), Lincoln Bryant (University of Chicago (US))
        • Rucio dump scripts in place for all endpoints
          • Chimera dump script for dCache
          • POSIX dump script for Ceph
          • Scripts running and dumps will be put into dCache today
        • New hardware status
          • UChicago
            • 18 Ceph Servers
            • Waiting on the bulk of the SSD and 10Gb cables
            • Will do a rolling online to rebalance as servers come online
          • Illinois
            • 20 C6320 servers (E5-2680, 256GB) - Adds 960 logical cores
            • HS06 is 11.53/core with HT enabled (48 logical cores/node)
            • HS06 is 553/node for a total addition of 11069
          • Indiana
            • 24 R630 (E5-250, 128GB) are now racked
            • Waiting on PDU to provide power
          • MWT2 accounting updated (v37, WLCG-v37, OIM, BDII, REBUS, Atlas Dashboard)
        • No Atlas jobs
          • Back filling with Opport jobs from OSG, Fermilab, Glow, LIGO, etc
          • Almost 11K cores
        • Indiana addition
          • Elizabeth Prout from OSG is now part time FTE for MWT2
          • She will handle the IU hardware and other issues
      • 14:30
        NET2 5m
        Speaker: Prof. Saul Youssef (Boston University (US))

        1) The Harvard gatekeeper had memory problems (now fixed) which caused several short interruptions on the Harvard side.

        2) We did an experiment moving the grid account home directories to an SSD, switching over when PanDA production dropped on Friday.  There was a brief DDM interruption because of this, but otherwise it was a smooth transition.  The solution is looking good so far.

        3) Finalizing FY15 purchases for 550TB of storage and 24 nodes (in C6300 enclosures) Intel Xeon E5-2660 v3 2.6GHz.

        4) Rucio consistency checking dumps have been provided as requested.

        5) We're in contact with Brian Bockelman and Brian Lin re: finishing up the transition to HTCondor-CE on the BU and HU sides (Harvard will also switch from LSF to SLURM).  

        6) Checking the Nov. WLCG accounting numbers.

        - Saul

      • 14:35
        SWT2-OU 5m
        Speaker: Dr Horst Severini (University of Oklahoma (US))

        - smooth operations, no issues

        - all OU and LU storage consistency dumps verified and tickets closed

        - network problems at LU were caused by switch problems, resolved

        - accounting numbers consistent and reasonable

         

      • 14:40
        SWT2-UTA 5m
        Speaker: Patrick Mcguigan (University of Texas at Arlington (US))

        UTA_SWT2

        • Power upgrade at UTA_SWT2 Facility was aborted.
          • Is rescheduled for 12/28/15
          • Will update OIM
        • New network path from Facility to Campus  (potentially more bandwidth)
        • Storage dump complete, needs to be automated in cron job

        SWT2_CPB

        • Power upgrade at SWT2_CPB is completed.
        • Storage dump complete, needs to be automated in cron job
        • Held Meeting with UTA-Networking concerning LHCOne and ScienceDMZ

         

      • 14:45
        WT2 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
        • HTCondor-CE operation is smooth. No longer see lost heartbeats in Panda (was due to the CE incorrectly marks running jobs as completed)
        • Installing the top-of-the-rack switch for the new blades. Next will be to test configuration with bare metal host, 8 core VMs, 16 core VM, etc. to understand the overhead. 
        • Event Service jobs seems to be running OK. Still  have the issue about preempting those jobs. May need to change the job definitions (Existing jobs run for many hours, and can't be preempted. So we can't use them to fill the opportunistic resources)
        • Power work today (Wednesday) and in January.
        • Working on spending the remaining 2015 funds on blades (same configuration)
    • 15:15 15:20
      AOB 5m