Speaker
Description
The growing complexity and scale of modern scientific computing infrastructures, such as the Port d’Informació Científica (PIC), a Tier-1 center within the Worldwide LHC Computing Grid (WLCG), require continuous optimization to maintain performance, reliability, and energy efficiency. Artificial Intelligence (AI) and Machine Learning (ML) techniques provide powerful means to tackle these challenges by enabling predictive and adaptive management of computational resources.
This contribution presents ongoing and prospective research directions focused on identifying and addressing operational challenges in large-scale computing environments through data-driven approaches. Illustrative examples include predicting the reduction in resource/cores utilization during compute-farm drainage periods, enhancing data cache management via intelligent eviction policies beyond traditional LRU mechanisms, identifying “hot” files to dynamically migrate them to higher-performance storage tiers (such as SSDs), and forecasting job execution times to improve scheduling and throughput.
Beyond the technical scope, this initiative aims to engage young researchers and students through short, well-defined projects that offer hands-on experience with real operational data and infrastructure from the PIC center. In addition, a set of future project ideas will be presented to inspire new collaborations and exploratory studies. By bridging educational efforts with practical challenges, this initiative seeks to foster innovation, attract emerging talent to the field of scientific computing, and contribute to the sustainable advancement of large-scale computing facilities within the WLCG ecosystem.
| Speaker release | Yes |
|---|