Conference on Computing in High Energy and Nuclear Physics

Name: Conference on Computing in High Energy and Nuclear Physics
Start: 2024-10-19T08:00:00+02:00
End: 2024-10-25T18:30:00+02:00
Location: No location set

19–25 Oct 2024

Europe/Zurich timezone

Contact Program Chairs

chep2024-pc@cern.ch

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

WED 02

23 Oct 2024, 15:18

57m

Exhibition Hall

Poster Track 4 - Distributed Computing Poster session

Marco Mascheroni (Univ. of California San Diego (US))

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters.

As part of our team's initiative to modernize infrastructure and enhance scalability, we undertook the migration of the glideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies.

Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Jeffrey Michael Dost (Univ. of California San Diego (US)) Marco Mascheroni (Univ. of California San Diego (US))

Brian Paul Bockelman (University of Wisconsin Madison (US)) Mr Colby Walsworth (University of California San Diego) Edita Kizinevic (CERN) Frank Wurthwein (UCSD) John Thiltges (University of Nebraska Lincoln (US)) Merina Albert (Fermi National Accelerator Lab. (US)) Vaiva Zokaite (Vilnius University (LT))

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory-6.pdf

Conference on Computing in High Energy and Nuclear Physics

Contact Program Chairs

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

Exhibition Hall

Speaker

Description

Authors

Co-authors

Presentation materials

Choose timezone

Conference on Computing in High Energy and Nuclear Physics

Contact Program Chairs

Speaker

Description

Authors

Co-authors

Presentation materials