Speaker
Description
We present operational case studies of real-world network incidents in the WLCG infrastructure, using all available metrics collected by perfSONAR toolkits - traceroute, latency, packet loss, and throughput - to diagnose when routing changes affect data movement. The paper focuses on how routing-aware triage helps distinguish impactful incidents from routine path variability.
The case studies include a router migration, a hardware-related reroute through commercial transit, a recurring undeclared route regime, and a live submarine cable fault detected before operator confirmation. Together, they show that combining path and performance measurements exposes failure modes that single-metric monitoring can miss, including lower-capacity detours, lost measurement visibility, and recurring structural deviations with measurable performance impact.
We show that per-corridor baselining can suppress non-impactful routing changes while surfacing a small set of corridors requiring operator attention. Looking toward HL-LHC operations, we argue that WLCG Data Challenges (DC) can benefit directly from this kind of routing-aware diagnostic framework: it can provide DC coordinators with a second source of operational signal, where routing anomalies are triaged and localized as they emerge rather than reconstructed hours or days later. This supports a broader view of HL-LHC readiness in which capacity planning is complemented by automated monitoring, rapid diagnosis, and shared incident records for validating future network operations.