US ATLAS Computing Integration and Operations

US/Eastern
Other Institutes

Other Institutes

Description
Notes and other material available in the US ATLAS Integration Program Twiki
    • 1
      Top of the Meeting
      Speakers: Michael Ernst, Robert William Gardner Jr (University of Chicago (US))
    • 2
      Production
      Speakers: Kaushik De (University of Texas at Arlington (US)), Mark Sosebee (University of Texas at Arlington (US))
      summary

      Kaushik:

      • Difficulty re-starting full production continue.
      • Causes: Validation? Prodsys?  Rucio? Software?
      • Tasks continue to be stuck for many reasons.  E.g. releases are not available in the US cloud - for no apparent reason.  Unknown error codes - thousands of tasks.  Jedi task buffer 100.
      • Commissioning chaos.  E.g. recon only running on high memory queues? 
      • New problems come with each release - and so production teams resist pulling the trigger, due to the plethora of software bugs. 
      • Not the production system itself - but that tasks are getting stuck.
      • What about user analysis?  Alden: not much traffic on the DAST list. Mostly physics analysis tools questions. Michael: analysis appears to be largely in tact - so at BNL, allow analysis to take over.
      • Kaushik: thefore the problem is likely limited to large scale production.

       

       

       

    • 3
      Data Management
      Speaker: Armen Vartapetian (University of Texas at Arlington (US))

       

       

    • 4
      Data transfers
      Speaker: Hironori Ito (Brookhaven National Laboratory (US))

      Hiro:

      • Very little traffic in entire ATLAS
      • 8 MB/s in aggregate!
    • 5
      Networks
      Speaker: Dr Shawn McKee (University of Michigan ATLAS Group)
    • 6
      FAX
      Speakers: Ilija Vukotic (University of Chicago (US)), Wei Yang (SLAC National Accelerator Laboratory (US))

      Ilija

      • Leaving sites untouched during the last two weeks.
      • There are some issues to address - called small meeting with Kaushik, Tadashi, and Paul.
      • So we now have more data to reconstruct what the algorithm.  
      • Notes difficulty in following what jobs are doing and why they might be failing.
      • During the break, US FAX infrastructure was very stable.

      Wei

      • EPEL repo refusing to adopt Xrootd 4.1, which has stability fixes. 
    • Site Reports
      • 7
        BNL
        Speaker: Michael Ernst

        Added 74 compute nodes - replacing aging equipment. Adds 3000 slots. 

        Hiro's group is preparing for a dCache upgrade.  2.10.. (not 2.11)

         

      • 8
        AGLT2
        Speakers: Robert Ball (University of Michigan (US)), Dr Shawn McKee (University of Michigan ATLAS Group)

        Bob:

        • MSU - 16 R620's to be installed.  Storage to come.
        • UM - equipment ordered, three more.
        • Implemented c-groups in Condor.  Easy but there are gotchas.  Does not play well with request memory parameters.
      • 9
        MWT2
        Speaker: Robert William Gardner Jr (University of Chicago (US))

        Equipment purchases

        CCC development - working with the Rucio team.

        dCache upgrade

      • 10
        NET2
        Speaker: Prof. Saul Youssef (Boston University (US))

        Purchase: storage purchase: 760 TB usable, MD series. 15 R630s.

        FTS performance, http://egg.bu.edu/atlas/adc/fts/plots/

        • All functional test transfers are failing worldwide.
        • #TCP streams to 1 for T1-T2 transfers set at BNL, better than floating?
        • Question about this.  

         

      • 11
        SWT2-OU
        Speaker: Dr Horst Severini (University of Oklahoma (US))

        All was smooth over the break.

        Quotes for R630s. 

        There will be a relocation of the move - to space with SciDMZ, 100 Gbps, fiber channel.  Timeframe unknown.  

      • 12
        SWT2-UTA
        Speaker: Patrick Mcguigan (University of Texas at Arlington (US))

        Patrick:

        • Smoothly during the break - missing dataset fouling up HC, Hiro taking care of it.
        • Bringing up 30xR620's.  Downtime likely required. Need newer version of the OS - Rocks update needed.  
        • Meeting tomorrow at UTA for next round of purchases. 
      • 13
        WT2
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))

        Wei

        Stable until yesterday.

        Power outtage tomorrow.

        Still working on process to order more storage.  Replacing Solaris - 4 head nodes, 8 trays of storage.  $120k. 

        Deploying 100 Gbps network between SLAC and ESnet.

    • 14
      AOB