US ATLAS Computing Integration and Operations

US/Eastern
Other Institutes

Other Institutes

Description
Notes and other material available in the US ATLAS Integration Program Twiki
    • 13:00 13:15
      Top of the Meeting 15m
      Speakers: Michael Ernst, Robert William Gardner Jr (University of Chicago (US))
    • 13:15 13:25
      Production 10m
      Speakers: Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US))
      summary

      Kaushik:

      • Difficulty re-starting full production continue.
      • Causes: Validation? Prodsys?  Rucio? Software?
      • Tasks continue to be stuck for many reasons.  E.g. releases are not available in the US cloud - for no apparent reason.  Unknown error codes - thousands of tasks.  Jedi task buffer 100.
      • Commissioning chaos.  E.g. recon only running on high memory queues? 
      • New problems come with each release - and so production teams resist pulling the trigger, due to the plethora of software bugs. 
      • Not the production system itself - but that tasks are getting stuck.
      • What about user analysis?  Alden: not much traffic on the DAST list. Mostly physics analysis tools questions. Michael: analysis appears to be largely in tact - so at BNL, allow analysis to take over.
      • Kaushik: thefore the problem is likely limited to large scale production.

       

       

       

    • 13:25 13:30
      Data Management 5m
      Speaker: Armen Vartapetian (University of Texas at Arlington (US))

       

       

    • 13:30 13:35
      Data transfers 5m
      Speaker: Hironori Ito (Brookhaven National Laboratory (US))

      Hiro:

      • Very little traffic in entire ATLAS
      • 8 MB/s in aggregate!
    • 13:35 13:40
      Networks 5m
      Speaker: Dr Shawn McKee (University of Michigan ATLAS Group)
    • 13:40 13:45
      FAX 5m
      Speakers: Ilija Vukotic (University of Chicago (US)), Wei Yang (SLAC National Accelerator Laboratory (US))

      Ilija

      • Leaving sites untouched during the last two weeks.
      • There are some issues to address - called small meeting with Kaushik, Tadashi, and Paul.
      • So we now have more data to reconstruct what the algorithm.  
      • Notes difficulty in following what jobs are doing and why they might be failing.
      • During the break, US FAX infrastructure was very stable.

      Wei

      • EPEL repo refusing to adopt Xrootd 4.1, which has stability fixes. 
    • 13:45 14:45
      Site Reports
      • 13:45
        BNL 5m
        Speaker: Michael Ernst

        Added 74 compute nodes - replacing aging equipment. Adds 3000 slots. 

        Hiro's group is preparing for a dCache upgrade.  2.10.. (not 2.11)

         

      • 13:50
        AGLT2 5m
        Speakers: Robert Ball (University of Michigan (US)), Dr Shawn McKee (University of Michigan ATLAS Group)

        Bob:

        • MSU - 16 R620's to be installed.  Storage to come.
        • UM - equipment ordered, three more.
        • Implemented c-groups in Condor.  Easy but there are gotchas.  Does not play well with request memory parameters.
      • 13:55
        MWT2 5m
        Speaker: Robert William Gardner Jr (University of Chicago (US))

        Equipment purchases

        CCC development - working with the Rucio team.

        dCache upgrade

      • 14:00
        NET2 5m
        Speaker: Prof. Saul Youssef (Boston University (US))

        Purchase: storage purchase: 760 TB usable, MD series. 15 R630s.

        FTS performance, http://egg.bu.edu/atlas/adc/fts/plots/

        • All functional test transfers are failing worldwide.
        • #TCP streams to 1 for T1-T2 transfers set at BNL, better than floating?
        • Question about this.  

         

      • 14:05
        SWT2-OU 5m
        Speaker: Dr Horst Severini (University of Oklahoma (US))

        All was smooth over the break.

        Quotes for R630s. 

        There will be a relocation of the move - to space with SciDMZ, 100 Gbps, fiber channel.  Timeframe unknown.  

      • 14:10
        SWT2-UTA 5m
        Speaker: Patrick Mcguigan (University of Texas at Arlington (US))

        Patrick:

        • Smoothly during the break - missing dataset fouling up HC, Hiro taking care of it.
        • Bringing up 30xR620's.  Downtime likely required. Need newer version of the OS - Rocks update needed.  
        • Meeting tomorrow at UTA for next round of purchases. 
      • 14:15
        WT2 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))

        Wei

        Stable until yesterday.

        Power outtage tomorrow.

        Still working on process to order more storage.  Replacing Solaris - 4 head nodes, 8 trays of storage.  $120k. 

        Deploying 100 Gbps network between SLAC and ESnet.

    • 14:45 14:50
      AOB 5m