US ATLAS Computing Integration and Operations

US/Eastern
virtual room (your office)

virtual room

your office

Description
Notes and other material available in the US ATLAS Integration Program Twiki
    • 13:00 13:15
      Top of the Meeting 15m
      Speakers: Michael Ernst, Robert William Gardner Jr (University of Chicago (US))
      • Revisiting Facility-wide deliverables from the US ATLAS Technical Planning Meeting 15m
        Speaker: Robert William Gardner Jr (University of Chicago (US))
    • 13:15 13:25
      Production 10m
      Speaker: Mark Sosebee (University of Texas at Arlington (US))
    • 13:25 13:30
      Data Management 5m
      Speaker: Armen Vartapetian (University of Texas at Arlington (US))
    • 13:30 13:35
      Data transfers 5m
      Speaker: Hironori Ito (Brookhaven National Laboratory (US))
    • 13:35 13:40
      Networks 5m
      Speaker: Dr Shawn McKee (University of Michigan ATLAS Group)
    • 13:40 13:45
      FAX 5m
      Speakers: Ilija Vukotic (University of Chicago (US)), Wei Yang (SLAC National Accelerator Laboratory (US))
    • 14:05 15:05
      Site Reports
      • 14:05
        BNL 5m
        Speaker: Michael Ernst

        Smooth operations in all aspects. Nothing specific to report.

      • 14:10
        AGLT2 5m
        Speakers: Robert Ball (University of Michigan (US)), Dr Shawn McKee (University of Michigan ATLAS Group)

        Twelve 48-core (HT) R630 servers are now on line at UM running HTCondor jobs.  These machines are dynamically configured within HTCondor.  Half have 128GB, half have 256GB, as the "as-delivered" machines with 192GB did not perform to their maximum capability.  HS06 measurements on these E5 2680 v3 processors with all 3 memory configurations (128GB and 256GB give identical results) are published at

        http://www.usatlas.bnl.gov/twiki/bin/view/Admins/HepSpecBenchmarks.html#SL6_results_32_bit
        and
        http://www.usatlas.bnl.gov/twiki/bin/view/Admins/HepSpecBenchmarks.html#SL6_results_64_bit

        A problem with the AGLT2 LMEM queue was fixed on Monday that sent pilots into the UNSUB state.  The problem was traced back to a configuration mistake in the APF queue.

        All AGLT2 gatekeepers were upgraded on Friday to OSG 3.2.31.  The update was live and was straight-forward.  This version includes HTCondor 8.2.10.

        Note that there is a problem in osg-configure, really a WARNING, that "max_wall_time" is not defined for each GIP sub-cluster, and so, the claim goes, it will default to 1440 for each.  The max_wall_time parameter is not actually described in the OSG documentation, and an OSG ticket (# 27580) has been opened on this deficiency.  I do not know if this parameter is actually used or propagated to anywhere important.

         

      • 14:15
        MWT2 5m
        Speakers: David Lesny (Univ. Illinois at Urbana-Champaign (US)), Lincoln Bryant (University of Chicago (US))
        • New hardware status
          • UChicago
            • 18 Ceph Servers racked and being tested
            • Will be configured for storage next week
          • Illinois
            • 20 6320 servers racked and being tested
            • E5-2680
            • 256GB, 2 x 1TB SSD
            • Will be released by ICC next week
          • Indiana
            • 24 R630 Servers on order
            • E5-2650
            • 128GB, 4 x 1TB disks
        • LIGO enabled and operation
        • CVMFS Server 2.1.20 complete
        • Upgrade to OSG 3.3.5, HTCondor 8.4.2 and HTCondorCE
        • DDLesny has CERN Certificate
          • Atlas VO
          • GUMS
          • OIM
          • AGIS
        • DATADISK filled up because Central Deletion was off 3 days
      • 14:20
        NET2 5m
        Speaker: Prof. Saul Youssef (Boston University (US))

        Smooth operations in the past two weeks with only minor problems.  Planning to make FY15 purchases next week.  Storage (1.1 PB), worker nodes, 2 NFS servers and networking gear.

      • 14:25
        SWT2-OU 5m
        Speaker: Dr Horst Severini (University of Oklahoma (US))

        - rucio consistency check scripts implemented at OU (OCHEP and OUHEP) and LU

        - waiting to be tested

        - LU currently unreachable because of network issue, seems to be at OneNet

        - Joel scheduled downtime, since it's not clear when this can be resolved, with the holiday coming up

        - OU all working fine, but OSCER is also turned off by downtime because it uses LU's SE

         

      • 14:30
        SWT2-UTA 5m
        Speaker: Patrick Mcguigan (University of Texas at Arlington (US))

        Rucio dumps in place for the auditing process.  Will automate for next month.

        We are expecting electrical work to be done at SWT2_CPB next week.

        UTA_SWT2 is going down this weekend due to facility electrical upgrades.

      • 14:35
        WT2 5m
        Speaker: Wei Yang (SLAC National Accelerator Laboratory (US))
        • Worked with Brian Bockelman on the HTCondor-CE issue (the CE prematurely marks some jobs as completed and deletes their  x509 proxies, etc. , and leave the LSF jobs running/wasting CPU). Brian has a hypothesis, and provided a patch. After the patch, we have't see this issue for 5 days.
        • 11 R630 blade batch nodes arrived. Waiting for top-of-the-rack switch. Estimated arrival of TOR is 12/8. Will test OpenStack on them first. Targeted full production date: Mid-Feb or End of Feb.
        • SMR based storage PO is in approval pipeline. ~1PB usable. 
        • Will use the remaining FY15 funds for batch nodes. Will start a discussion with the SLAC computing center about Dell's equipment leasing model.
        • some electric works coming
    • 15:05 15:10
      AOB 5m