28 September 2026 to 2 October 2026
Castelldefels, Barcelona, Spain
Europe/Zurich timezone

Characterization of Reset-to-Reset Latency Stability in the CBM Timing and Fast Control System

29 Sept 2026, 13:40
1h 40m
Castelldefels, Barcelona, Spain

Castelldefels, Barcelona, Spain

Hotel Rey Don Jaime
Poster Timing & Trigger Distribution Poster 1

Speaker

Vladimir Sidorenko (Karlsruher Institut für Technologie)

Description

Accurate clock and time distribution is a key requirement for the free-streaming data acquisition of the CBM experiment, where timestamped data are reconstructed online at interaction rates of up to 10 MHz. The Timing and Fast Control system distributes the global clock and time to about 200 readout FPGA boards over a hierarchical network of latency-deterministic optical links. A previous study addressed in-run latency variation and its accumulation over multiple hops. The presented work extends this characterization to reset-to-reset latency stability, which determines the reproducibility of time distribution after system reinitialization and is critical for scaling to the final experiment.

Summary (500 words)

The Compressed Baryonic Matter (CBM) experiment is being developed to study strongly interacting matter at high baryonic densities at interaction rates of up to 10 MHz. Due to the complexity of the relevant trigger signatures, event selection cannot be efficiently implemented with a conventional hardware trigger. CBM therefore uses a free-streaming data acquisition architecture, where the self-triggered front-end electronics continuously timestamp and transmit detector data to a computing farm. There, the First-Level Event Selector performs time-based event reconstruction and selection completely online. With up to 1 TB/s of timestamped data transported through the readout chain, the quality of the distributed time reference directly affects the quality of event reconstruction.

The common clock and time are distributed by the Timing and Fast Control (TFC) system. In the CBM readout chain, detector data are concentrated by GBTx-based readout electronics and transported over 4.8 Gb/s optical GBT links to approximately 200 Common Readout Interface (CRI) FPGA boards. These CRI boards form the entry stage of the online computing infrastructure and also provide the timing reference further downstream to the front-end electronics. The TFC system must therefore synchronize the CRI layer and provide a stable time base to the experiment. In addition, it provides a low-latency path for fast-control information, such as status messages and throttling commands used to protect the data acquisition system against congestion caused by beam-intensity fluctuations.

To serve the required number of endpoints, the TFC system is implemented as a hierarchical FPGA network with Master, Submaster and Endpoint nodes connected by bidirectional optical links. The Master node defines the global time reference and distributes clock and time messages downstream through one or more Submaster layers to the CRI endpoints. The clock is recovered from the incoming optical link at each node and reused for further downstream transmission. This approach provides a scalable architecture, but it also makes deterministic link behavior a critical requirement, since variations of the downstream latency directly translate into timing offsets between endpoints.

Two types of latency variation are relevant for such a system. In-run variation describes the short-term latency fluctuations during uninterrupted operation and is closely related to jitter. However, correct operation of the final experiment also requires reproducible timing after power cycles, FPGA reconfiguration, or complete system restarts. This is governed by reset-to-reset latency stability, which can originate from transceiver initialization, clock-domain alignment, link bring-up procedures, and deterministic phase reconstruction in cascaded optical paths.

The present contribution focuses on this reset-to-reset aspect of the TFC data transport. It investigates how reproducibly the downstream latency is established after repeated reinitialization of the system and how this behavior scales in a hierarchical network. For this purpose, an automated characterization procedure is used to repeat link and system initialization, measure the resulting message transport latency, and separate reset-to-reset effects from in-run jitter. The study will discuss the measurement concept, the relevant latency-critical parts of the TFC architecture, and the implications of reset-to-reset stability for the final CBM timing distribution network.

Author

Vladimir Sidorenko (Karlsruher Institut für Technologie)

Co-authors

Felix Frombach (Karlsruher Institut für Technologie) David Emschermann (GSI Helmholtzzentrum fuer Schwerionenforschung GmbH) Dr Wojciech Zabolotny (Warsaw University of Technology (PL)) Walter Mueller (GSI Helmholtzzentrum fuer Schwerionenforschung GmbH) Juergen Becker (Karlsruher Institut für Technologie)

Presentation materials