31 August 2026 to 4 September 2026
US/Pacific timezone
All in-person registration fee waivers have now been claimed.

AI Applications Porting, Running and Scalability Analysis on Voyager – an AI-focused Intel/Habana Gaudi processor-based supercomputer

Not scheduled
20m
Tutorial Tutorials

Speakers

Dr Javier Hernandez Nicolau (San Diego Supercomputer Center, UC San Diego)Dr Madhusudan Gujral (San Diego Supercomputer Center, UC San Diego)Dr Mahidhar Tatineni (San Diego Supercomputer Center, UC San Diego)

Description

This tutorial will present the architecture, Kubernetes based systems setup, user software environment, scalability studies, and fine tuning on the Voyager system. Voyager is an US National Science Foundation funded AI-focused hardware based supercomputer. It is built using the Intel/Habana Gaudi processors (Gaudi1 and Gaudi2), has a 400 GbE interconnect from Arista for scale out training and provides a Ceph file system for storage. Gaudi AI processors natively integrate 10 (for Gaudi1) and 24 (for Gaudi2) 100-Gigabit Ethernet ports of RoCE v2 (RDMA over Converged Ethernet) on-chip, enabling flexibility of scaling and avoidance of throughput bottlenecks that can limit scaling capacity. Voyager’s AI-focused hardware and the associate system and user software environments are optimized specifically for Deep Learning (DL) AI applications that primarily use the PyTorch machine learning framework for AI applications. Scientific applications from various areas such as high energy physics, astrophysics, radiation oncology, psychology, biomedical text analytics and neuroscience have been ported and scaled on the Gaudi architectures of the Voyager machine. LLM based fine tuning has been extensively studied on this machine for various data sets. Voyager is available for research and education to the scientific user community via the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) and the National AI Research Resource (NAIRR) Pilot allocation process.
This tutorial will go over the Voyager architecture in detail, explain the Kubernetes based user environment, and show how to port and run applications on the Voyager system. We will cover how to run a Deep Learning (DL) model, based on the PyTorch framework, on the Voyager supercomputer.

As the first example, we will take a PyTorch application and will port it using Intel/Habana Gaudi libraries. We will discuss how to run Jupyter notebooks on Voyager. We will show how to pull a Hugging Face model and run inference on Voyager using Transformers/Diffusers and Optimum-Habana libraries.

As another example, we will show fine-tuning work on the Voyager machine. LLMs are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. We will focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration.

Tutorial level (only for Tutorial) Intermediate
Do you plan to submit a 4-page extended abstract on OpenReview (only for Presentations/Posters)? No

Author

Co-authors

Dr Javier Hernandez Nicolau (San Diego Supercomputer Center, UC San Diego) Dr Madhusudan Gujral (San Diego Supercomputer Center, UC San Diego) Dr Mahidhar Tatineni (San Diego Supercomputer Center, UC San Diego)

Presentation materials

There are no materials yet.