WLCG Sustainability Forum Meeting #6: Perspective from gravitational waves and astrophysics

→ Europe/Zurich
ZOOM (CERN)

ZOOM

CERN

Description

After surveying the landscape for HEP experiments (the main target for the WLCG strategy) in the previous meetings, we are now organising a meeting to understand the environmental sustainability strategies and needs of future experiments in gravitational waves and astrophysics that are building their distributed computing models - the Einstein Telescope and the Square Kilometer Array.

This meeting will include a short presentation from HEP/WLCG to set the scene, and two presentations, one for ET and one for SKAO. 

This meeting is intended as a first step to understand points of contacts and potential shared strategies and work.  

Zoom Meeting ID
69284088962
Host
Markus Schulz
Useful links
Join via phone
Zoom URL

Summary

1. Introduction to the WLCG Environmental Sustainability Forum (Caterina Doglioni)

  • The Worldwide Large Hadron Collider (LHC) Computing Grid (WLCG) is now 20 years old. Physics ambitions exceed what the available computing resources allow, so more physics must be done with the same or fewer resources while making computing more environmentally sustainable.
  • The Forum aims to discuss and act on the three WLCG strategy recommendations, which should be completed, or at least reviewed, next year:
    • Agree on metrics and a framework to collect energy-efficiency information.
    • Facilitate the use of more energy-efficient hardware, as a co-design effort across the many experiments using WLCG.
    • Develop and promote a sustainability plan covering the full lifecycle: software, computing models, facilities and hardware.
  • Progress so far:
    • A series of meetings since October last year, on power accounting and the embodied carbon of graphics processing units (GPUs).
    • Good progress on metrics.
    • Energy efficiency is being improved, mostly by moving to accelerators.
    • The sustainability plan must look ahead to the High-Luminosity LHC (HL-LHC), since computing-model decisions are being made now.
  • The Forum has focused on simulation and event generation because they use a large share of compute cycles, which makes software important.
  • Today's session looks outward, at what WLCG, the Square Kilometre Array (SKA) and the Einstein Telescope (ET) can do together. A further meeting on lessons learned is planned.
  • The European Union (EU) ENSURE project (Einstein Telescope, WLCG, SKA and others) aims to reduce the environmental footprint of large European research infrastructures in computing. It treats sustainability as a managed process: measure, model, then mitigate. Expected outputs are open-access tools and possibly environmental-impact-aware scheduling.
  • Participants are encouraged to start thinking about how to align strategies through ENSURE.

2. SKA and the SIMS Project (Denisa-Andreea Constantinescu, EPFL)

Project overview

  • SIMS (sustainable computing for the Square Kilometre Array, SKA) is a consortium project of École Polytechnique Fédérale de Lausanne (EPFL) groups and French partners (Institut National des Sciences Appliquées (INSA), the Observatory of Paris, the Observatoire de la Côte d'Azur). It is funded by the Swiss National Science Foundation (SNSF) and focused on the SKA Science Data Processor (SDP).
  • It aims to improve energy efficiency and sustainability through domain-specific computing platforms and energy-aware scheduling that minimises emissions. Demonstrators are planned by the end of the project, using real SKA data where possible.
  • Key constraints:
    • Tight power caps at both SKA sites, with essentially no spare capacity in Western Australia.
    • Limited budget, so the first iteration will use central processing units (CPUs).
    • Unstable pipelines, which favour flexibility over efficiency.
    • High grid carbon intensity in South Africa and Western Australia.
  • The SDP takes correlated visibilities from the Central Signal Processor (CSP) and produces images. The SIMS scope is the optimisation of this processing.

Results presented

  • CEO-DC framework (preprint on arXiv at https://arxiv.org/html/2507.08923, to be released as open source):
    • Models embodied and operational emissions, cost, power usage effectiveness (PUE) and hardware lifetime for procurement planning.
    • Carbon-tax analysis: the highest current tax is around $170 per tonne, while real GPU platforms would need roughly $600 per tonne to make sustainable options economically viable.
  • AstroCamp co-design framework for radio astronomy at https://arxiv.org/html/2512.13591v3:
    • Unified metrics covering performance, energy, system-level behaviour and economics.
    • Benchmarked on an EPFL Kuma node (listed in the Green500, the ranking of the most energy-efficient supercomputers).
    • Finding: larger image sizes do not give much higher throughput than smaller ones.
    • An open gap remains: astronomers still need to agree on accuracy and real-time requirements.
  • Imaging pipelines (WSClean and Image-Domain Gridding (IDG)):
    • They do not scale on multi-core CPUs or GPUs.
    • Utilisation is only 5–15%.
    • About 85% of the energy is static or idle, mostly because of input/output (I/O) bottlenecks, not because of the hardware itself.
    • The code is old and unoptimised.
    • This justifies custom domain-specific accelerators, but software and data formats must be improved first.

Discussion

  • The embodied-carbon formula in the talk is simplified. CEO-DC treats hardware lifetime and PUE as parameters (PUE of 1.2 in the example). The SKA would like a lifetime of at least 6 years, against a typical 4 for public institutions.
  • Markus Schulz noted that the optimal upgrade cycle depends on grid carbon intensity. France is roughly 25–30 grams of carbon dioxide (CO₂) per kilowatt-hour (gCO₂/kWh), while South Africa is about 672 gCO₂/kWh. Denisa noted that for artificial intelligence (AI) workloads the framework suggests upgrading about every 2 years. CERN (the European Organization for Nuclear Research) has lifecycle models that can be shared.
  • Markus noted that WLCG sites have high utilisation, so further efficiency gains there are mainly in software. Denisa's view was that software can give tens to hundreds of times improvement, scheduling about two orders of magnitude, and hardware at most about 2×.
  • Next steps for SIMS:
    • Adopt new parallel data-access technologies.
    • Use power caps and dynamic voltage and frequency scaling to cut idle energy.
    • Build accelerators on field-programmable gate arrays (FPGAs), moving towards a reconfigurable application-specific integrated circuit (ASIC) design.
    • Make scheduling more aware of system-level utilisation.
  • Planned ASIC scope: the kernels between gridding and the fast Fourier transform (FFT) (five or six kernels, including non-uniform FFTs), plus the CLEAN algorithm primitives.

3. Einstein Telescope Computing and Sustainability (Luciano Gaido, Istituto Nazionale di Fisica Nucleare, INFN)

Background and challenges

  • The Einstein Telescope (ET) is a proposed European third-generation gravitational-wave detector, working in a network with Cosmic Explorer. It is expected to start in the late 2030s or early 2040s. The site (three candidates) and the design (triangular or two L-shaped interferometers) are still to be decided.
  • It will have about 10 times greater sensitivity, a much higher event rate (about 1,000 times more events) and long, overlapping signals. This calls for new algorithms, longer templates and more random-access memory (RAM).
  • Computing target: stay at about 10% of a High-Luminosity Large Hadron Collider (HL-LHC) experiment (Virgo today is about 10% of an LHC experiment). This requires a 10–100× speed-up.
  • Data: strain data of 10–100 terabytes (TB) per year, and about 100 petabytes (PB) per year of raw data (mainly for archiving). All figures are overestimates and will be refined.

Computing model

  • Delivered to the European Commission under the Einstein Telescope Preparatory Phase project (ETPP), which ends this year and will be followed by the two-year ETCompass project.
  • Based on the LIGO-Virgo-KAGRA (LVK) model, where LIGO is the Laser Interferometer Gravitational-Wave Observatory and KAGRA is the Kamioka Gravitational Wave Detector. It uses Worldwide LHC Computing Grid (WLCG) tools for distributed computing where possible (e.g. Rucio, the High-Energy Physics (HEP) Software Foundation (HSF) Conditions Database, REANA), plus Open Science Grid (OSG) tools such as HTCondor and Pelican.
  • A new domain, rapid alerts:
    • The low-frequency cutoff (from 1–2 hertz) allows early warning hours or days before a merger.
    • Current low-latency processing takes about 30 seconds on 20,000 cores and about 500 GPUs.
    • Simply scaling this up would need about 40 million cores, which is impossible. Mitigations are artificial intelligence and machine learning (AI/ML), new GPU architectures and other techniques.

Sustainability

  • Energy estimates in the computing model use the WLCG methodology, with different power usage effectiveness (PUE) values and CPU-only versus mixed CPU/GPU scenarios. Lower PUE and mixed GPU/CPU setups are best.
  • Resources must be sustainable over a lifespan of about 50 years.
  • Needs identified: a common software framework, best practices, end-to-end testing, analysis platforms that avoid duplication, and monitoring and accounting to predict needs.
  • Related activities:
    • A green data-centre prototype in Germany, in a shipping container, with computing power adapting to energy availability.
    • ENSURE, connected through INFN.
    • M2 Tech, a project in preparation.
    • Community building and training.
  • Collaboration with WLCG, HSF, the Cubic Kilometre Neutrino Telescope (KM3NeT), the Cherenkov Telescope Array (CTA), the Square Kilometre Array (SKA) and others is seen as mandatory.

Discussion

  • LISA collaboration (Markus Schulz, Paul Laycock, Steven Schramm):
    • The Laser Interferometer Space Antenna (LISA) is the planned space-borne gravitational-wave detector. There is currently no direct data-analysis collaboration with it, and the frequency range is very different.
    • LISA signals last indefinitely, so it uses global fits with iterative source subtraction. LVK and ET use matched filtering.
    • Subtraction errors grow quickly at ET's precision, so ET's intermediate regime may need a hybrid approach. This is the biggest open data-analysis question for ET.
    • One area of interest is tracking objects handed over from LISA to ET.
  • Common ground with SKA:
    • Fast Fourier transforms (FFTs) are common to both projects.
    • Both are interested in application-specific integrated circuits (ASICs) for FFTs and non-uniform FFTs.
  • Numerical precision:
    • Virgo is currently discussing float32 versus float64, meaning 32-bit versus 64-bit floating-point numbers.
    • Strain values of about 10⁻²¹ approach float32 limits. Float32 is likely sufficient if the data is stored and offset more carefully, e.g. via a translation layer.
    • Lower precision (16-bit or float8) may work for ML-based calculations, while raw data stays in float32.
    • SKA is investigating custom bit widths, including block floating point, once real data is available.
    • Markus Schulz noted that ASICs need not follow the Institute of Electrical and Electronics Engineers (IEEE) floating-point formats and that bit widths and exponent ranges can be adapted.


Full transcript 

Part 1: Introduction to the WLCG Environmental Sustainability Forum (14:35–14:43)

Caterina Doglioni: I'll give a very brief introduction, mostly for the people here from the other experiments and projects. This is the obligatory introduction. WLCG is the Worldwide LHC Computing Grid. It is now 20 years old, and it handles a very large amount of computing. The slide shows the difference between then and now. I'm worried I won't quote the right numbers, but the message of the plots is what matters. I apologise for the bad audio. I had to attend a UK Parliament session earlier and couldn't get out of it.

The point here is that we have more ambition in physics than our resources allow, which is an intrinsic problem. There is a convergence of interests around energy awareness. We need to do more physics with the same or fewer computing resources, and we also want to make computation more environmentally sustainable. This is the underlying concept of the WLCG Environmental Sustainability Forum. Its goal is to discuss and act on the three recommendations of the WLCG strategy for 2027. We need to get these done, or at least take stock of them, next year.

The first recommendation is to agree on metrics and provide a framework to collect information related to energy efficiency. The second is to facilitate the use of more energy-efficient hardware wherever possible. This has to take into account that many experiments use WLCG, so it is a co-design problem rather than something we can impose on the experiments. The third is to develop and promote a sustainability plan to improve energy efficiency and reduce carbon footprint. This covers the whole lifecycle of LHC data processing: software, computing models, facilities, hardware and so on.

So far, we have held a series of meetings, starting in October last year, on power accounting and the embodied carbon of GPUs. That work was inward-looking. Today we are looking outward, to understand what WLCG, SKA and Einstein Telescope can do together. We will also hold a further meeting on experiences and lessons learned. We focus on this topic because simulation and event generation take up a lot of compute cycles, and software is important.

How far along are we? On the first recommendation we are making good progress, and there is also a European project that I will show on a slide. On the second, energy efficiency is being studied and improved, mostly by moving to accelerators. On the third, the sustainability plan, we have to look ahead. The HL-LHC starts in a few years, and decisions on the computing models are being made now.

There is also the ENSURE project, an EU project involving the Einstein Telescope, WLCG and SKA as partners. Its motivation is to reduce the environmental footprint of large European research infrastructures in computing. The project approaches environmental sustainability as a managed process. We need to measure and model, and then mitigate. The goals and outcomes are a suite of open-access tools and, potentially, green-aware scheduling, though I'm not sure we will use that. The aim is to align European digital science with climate goals. Since we are all part of this project, perhaps in this meeting we can start thinking about how to align our strategies. We will talk about the project itself separately, but today is a good introduction.

Are there any questions? I'm not seeing any, so I'll stop sharing and we can move on to the next talk, on the SKA, by Denisa-Andreea Constantinescu.


Part 2: SKA and the SIMS Project, Denisa-Andreea Constantinescu (14:43–15:13)

Denisa-Andreea Constantinescu: I'm Denisa-Andreea Constantinescu, a scientist at EPFL working on sustainable computing. I'm also the lead scientist and project manager of the SIMS project, which stands for sustainable computing for the SKA. We focus specifically on the Science Data Processor.

This is a consortium project. The partners in Switzerland are at EPFL: the Embedded Systems Laboratory, the Cloud Sustainability Center [name unclear] and SCITAS, the HPC support group at our university. We also have partners in France, including INRIA, the Observatory of Paris and the Observatoire de la Côte d'Azur. Of course, everyone working on the SKA Observatory science and the Science Data Processor is involved as well, including MeerKAT and others.

The goals of this project align well with this session: how to improve the energy efficiency and sustainability of computing. We aim to do this in two ways. The first is designing domain-specific computing platforms that are as efficient as possible. The second is scheduling, so that you co-optimise for energy, not just for performance, with the overarching goal of reducing emissions. By the end of the project we want to have demonstrators for both. I hope most of them will use real SKA data. However, the telescope is still being built. It will have very good specifications if it is funded all the way to the desired number of antennas, in both Western Australia and South Africa. I'm not sure we will manage to have demonstrators in the loop with the actual development of the Science Data Processor. We may use preliminary data from some of the intermediate milestones of the telescope.

Besides wanting to do as much science as possible with these new computing and scheduling techniques, and keeping latency as close as possible to real time, our biggest challenge is a cap on the power that the Science Data Processors can use at the two locations. Initially the cap was supposed to be 2 [unit and figures unclear], and it is now being reduced to 5. In Western Australia there is actually no capacity even for a few hundred kilowatts. This is a huge challenge, so we try to make things as efficient as possible.

To narrow the scope a little, the SKA will have several computing components. The Central Signal Processors (CSPs) are exactly where the data is collected at the two sites. Then there are the SKA Regional Centres across the globe, a little like what CERN is doing. We have a Science Data Processor. We take correlated visibilities from the CSP, optimise the execution of these pipelines, and produce images as output. Scientists can then use these images for their science.

Why do we care so much about sustainability? I'll give you the platform-design angle. Data centres in general, not just scientific ones, are going to consume more and more energy. Energy is expensive, and it is also very carbon-intensive. The energy mix is not great in South Africa or Australia, so the less energy we consume to operate the platforms, the better. The newest commercial platforms, CPUs and GPUs, are also the most performant, but they carry more and more embodied carbon from their production. So our question is whether we can optimise the design process to minimise both embodied and operational emissions.

There is a third dimension for the SKA: cost. Computing budgets are limited. For the first iteration, for example, they can only afford CPUs. The pipelines are also still unstable, so you want something very flexible, but that is not very energy-efficient. If utilisation is poor, you have to over-procure to meet the compute demand, so the cost ends up much higher than it should be. We have to juggle all of these things at once.

On hardware design, there are two things we can look at. One is energy efficiency, which directly affects operational emissions. The other is the acceleration you can achieve with a hardware accelerator. If you can go two times faster at the same platform cost, you need to procure only half the units for the same computation, which reduces the embodied emissions of that procurement. Ideally we also keep the cost as small as possible. So the question is how much you need to improve energy efficiency and acceleration to reduce net emissions and perhaps also save money, compared with the commercially available platforms. That is the target when we make design decisions.

There are different types of platforms: CPUs, GPUs, FPGAs and ASICs. ASICs are quite expensive, and it is difficult to set up a production line for one very specific chip design, but you get the best energy efficiency. CPUs are the opposite. When you quantify the emissions and cost of each, there is a clear trade-off. Location matters too. Australia has a somewhat better carbon intensity than South Africa, so the decision on which platform to use at each site might depend on this.

I'll briefly present some results from the SIMS project. One is the CODC framework, which is already on arXiv as a preprint. It models all the factors that a data-centre architect, or a hardware or software designer for a specific domain, has to take into account. You optimise acceleration and energy efficiency to reach the sweet spot where being sustainable is also economically viable. Sometimes it is impossible to get there, and we realised that we would need incentives such as carbon taxes. Right now, I think the highest carbon tax in the world is around 170-something dollars per tonne. When we crunched the numbers for real GPU platforms, we found this figure should be around 600. That is quite expensive, so we push hard to make better platforms rather than make people pay more for carbon, which would be a very unpopular outcome. We aim to make this framework open source, so that people can use it for procurement planning: to choose which platforms to mix to reach a specific computational throughput while minimising lifecycle carbon emissions. I may brief you on this another time, when everything is online.

The other outcome of the SIMS project is AstroCamp. When I started working on this project, I realised there were no clear metrics for the scientific and computing aspects: how much throughput, and how much precision and accuracy, your pipelines need. These pipelines are still work in progress. They are not even working yet, although there are some antennas in the field. Our project is not directly involved in SKA decision-making. It is an external project funded by the Swiss National Science Foundation (SNSF). We have one person in the SKA co-design group who feeds us information about the official procurement and co-design. But it is quite hard to get an astronomer to say what accuracy and real-time constraints they need. So we decided to build a flexible co-design framework where you can parameterise everything and guide some decisions early on, without committing to anything.

We included a unified set of metrics. There are the classical ones for platforms, such as performance and energy. System-level metrics are also quite important when you scale from a single GPU to a node, a rack or a whole cluster. Then there are the economics. The missing box is the one we need to agree on with the astronomers, which could be an extension in a future release of AstroCamp.

Here you can see some examples of benchmarking on a Kuma node at EPFL, which is now listed in the Green500. One interesting finding is that even with the largest image output in our benchmark, you don't get much higher throughput than with smaller image sizes. This was interesting for the astronomers, who thought they would have to stick with smaller image sizes to maximise throughput and latency. It is an insight you can get just by playing with this framework.

To wrap up, here are some key insights on imaging pipelines. Examples are WSClean and IDG, which LOFAR currently uses and which will also be used in the initial phase of SKA-Low. We saw that the code doesn't scale at all on multi-core CPUs, multiple CPUs or GPUs. Utilisation is very poor. Static energy is most of the energy consumed, and the dynamic energy, the useful work, is really poor. After a couple of years of work, this justifies building custom domain-specific accelerators for these workloads, because what is on the market is not the best solution. The code is also not very well optimised, and much of it dates from decades ago. Some of it is better parallelised and maintained. So a co-design effort is needed to reach the targets that the SKA has. That's it from my side. Let me know if you have any questions. I hope I didn't overrun. It was supposed to be under 15 minutes, or 10, I don't remember.

Caterina Doglioni: Thank you very much. Are there any questions from the Zoom room?

Denisa-Andreea Constantinescu: Markus, please go ahead.

Markus Schulz: You have a slide with embodied CO2 and the formula for how you calculate it. You have power times time times carbon intensity, but you also have to account for the embodied CO2 over some lifetime. Here it seems to be a static number.

Denisa-Andreea Constantinescu: It is simplified, indeed. When we do the analysis, we take into account the lifetime of the procurement. In CODC, for instance, it is a parameter you can choose. There are also other elements, such as power usage effectiveness (PUE). We don't take water usage effectiveness into account, but the ratio of cooling efficiency to compute is also important. For simplicity, I didn't want to go through all the parameters.

Markus Schulz: So you run your studies with a fixed PUE, not depending on the site?

Denisa-Andreea Constantinescu: In this publication, we set it as a parameter. In this example we used 1.2, but it is flexible. If you know the PUE of your data centre, you can use that.

Markus Schulz: 1.2 is pretty good. What do you estimate as the lifetime of your hardware?

Denisa-Andreea Constantinescu: Given that the SKA budget is not amazing, they would like to keep their platforms for at least six years. When a university or public institution purchases hardware, the usual time is four years, but I think they would like to keep it for longer.

Markus Schulz: It's a complicated game, because if you buy more often, you get more performance for the same power. We have people working on models for the overall lifecycle. We can send you a link if you're interested.

Denisa-Andreea Constantinescu: Yes, we did a study of this for AI computing with the same framework. In many cases, for the most modern models, you would want to upgrade every two years or so. The improvement in both energy efficiency and performance is so great that by not upgrading more often you would lose money on operation, and you would lose the computing capacity you want to provide as a service.

Markus Schulz: It also depends on the CO2 intensity of your electricity. In France, it ranges from 25 to 30-something grams per kilowatt-hour. In other places it can be 700. So for both cases you have different optimisations.

Denisa-Andreea Constantinescu: Yes, it's one more parameter, because we have to optimise for each location. South Africa and Western Australia both have a poor carbon intensity. I don't know the number by heart.

Markus Schulz: It's 672 grams per kilowatt-hour.

Denisa-Andreea Constantinescu: Yes. Any more questions?

Caterina Doglioni: I have one on the slide where static energy is around 85% of the budget. Is this idle power? Is it specific to your workload, or more general?

Denisa-Andreea Constantinescu: I didn't fully understand the question because of the noise, but I understand you're asking about idle power. This is specific to the workload. The deployment is not ideal. The utilisation of the machines is between 5% and 15%. Most of the cores and most of the memory are not utilised properly, and it doesn't scale beyond one platform. So I wouldn't say this applies to every kind of workload, and I believe these pipelines can be optimised for better utilisation. When you zoom into the whole pipeline, you see some kernels, for instance FFTs or gridding, that are super-optimised. Gridding on the GPU reached 80% utilisation, so the dynamic power would be much higher and the idle share much lower. But there are so many blocking points between the different components of the pipeline that most of the time the hardware is starved and idle. It consumes energy while idle, because there is always a baseline static power. Markus, I hope this answered your question, Caterina.

Markus Schulz: I don't know whether Caterina is happy, but I'm still confused. You say it's blocked. Does this mean you go parallel and then need to aggregate the results, or is it I/O that starves you?

Denisa-Andreea Constantinescu: It's mostly I/O. One problem I've seen is the data format. It is formatted in a way that humans can understand, but it is not optimised for parallel access, nor for the data representation or compression that the data actually needs. The visibility files are tens of gigabytes, or even a terabyte. At some point a bunch of threads all want to read or write the same file. The way these pipelines are written, I think there was no clear thought given to parallel data access and data flow. The work we've been doing is not on optimising the software itself. We are building hardware that improves energy efficiency and acceleration, and scheduling that improves system-level efficiency. But after this first iteration of benchmarks, it's clear that the software has to be improved first, and the data formats and structures as well.

Markus Schulz: So you have a huge room for improvement in the software. But this is true in many places. We don't have the static-energy problem, because we have very high utilisation rates. We can even raise effective rates by running more jobs than we have cores on the same platform. The improvements in power efficiency that we can still get are mainly in software, and this is really a challenge.

Denisa-Andreea Constantinescu: Yes. You can get tens or hundreds of times improvement in energy efficiency by improving your software, but from hardware you cannot hope for more than 2x. At a higher level of abstraction, scheduling can give you something in between: about two orders of magnitude. Any more questions? I see something in the chat: what are the plans for the SIMS project going forward, given what we learned?

One lesson concerns the data format, and we're trying to get rid of that problem. Some people are working on improving parallel access, and we will adopt any new technology in our experimental pipelines. The other direction is working with power caps and dynamic voltage and frequency scaling of the platforms. Where you cannot improve utilisation at all, you can at least lower as much as possible the static and idle energy you are burning while doing nothing. A third direction: I didn't present everything, as I only had about 15 minutes. We are working heavily on building domain-specific accelerators using FPGAs. We are also moving towards an ASIC design, which would be the most energy-efficient option, but with reconfigurability. We know the software is still changing. They haven't set specifications for the desired image resolution or the number of channels for the imaging pipelines. So we are trying to get as close as possible to the most energy-efficient design, which is an ASIC, while also making the scheduling more aware of system-level utilisation, if utilisation cannot be improved otherwise. Yes, Markus?

Markus Schulz: Can you give an example of what kind of processing you plan to do with ASICs? We have ASICs inside the detectors, but we have not considered them outside the detectors.

Denisa-Andreea Constantinescu: This is a very coarse-grained representation of the kind of computing we have in the Science Data Processor. We are thinking of putting two pieces together. Do you know about non-uniform FFTs?

Markus Schulz: No.

Denisa-Andreea Constantinescu: When you go from the frequency domain to the image domain, you can use an FFT. In this case, however, the sampling of the data is not uniform. It is somewhat randomly spaced, depending on the data from the antennas and the configuration. So you need a gridding operation, which is a bit like a convolution, but multi-dimensional. We are thinking of putting all the necessary kernels between the gridding and the FFT stages on an ASIC. There are five or six different kernels. Inside this block is CLEAN, the most famous algorithm in radio astronomy. It uses very basic, recurrent functions, such as finding the maximum intensity, max and argmax, and sine and cosine functions of the kind used inside FFTs. These operations can be very well optimised. We will also take into account the integration between the different stages of the pipeline, because there is an order to the operations. Sometimes you have to transpose the data or change the way you structure it to do the computation. So we are thinking about how to do that on the fly, in a more efficient way.

Markus Schulz: Thanks.

Denisa-Andreea Constantinescu: My pleasure. Any more questions? Okay, that's it from my side. I'm going to stop sharing now.

Caterina Doglioni: Thank you very much, perfect. There's also a question from the chat. We can now move on to the Einstein Telescope presentation.


Part 3: Einstein Telescope Computing and Sustainability, Luciano Gaido 

Luciano Gaido (INFN): This presentation was prepared with contributions from several people, some of whom are already attending this meeting. I will start with a brief introduction to gravitational waves, and then talk about the sustainability activities within the Einstein Telescope (ET) community.

Gravitational waves were theoretically predicted by Einstein in 1915 and experimentally detected by LIGO in 2015. To detect them, we use a modified Michelson interferometer. What is really important to measure is the strain, the differential arm length of the interferometer, which is sampled at 16 kHz.

Today we have the so-called second-generation detectors. There are two LIGO detectors in the US, and a third is being built in India. In Europe we have two: Virgo in Italy and GEO600 in Germany, which is being phased out this year. Then there is KAGRA in Japan. LIGO, Virgo and KAGRA are closely connected through a cooperation called IGWN, the International Gravitational-Wave Observatory Network, as well as the LVK collaboration. They share the computing infrastructure and the data, which is really important for scalability and also for multi-messenger physics.

Among third-generation detectors, the Einstein Telescope is the proposed European one. ET will work in a network with another proposed detector, Cosmic Explorer. ET will be a more complex detector, including a low-frequency interferometer in addition to the standard high-frequency one, to reach a sensitivity a factor of 10 greater. This means a much higher information density, of the order of [figure unclear in the recording] more. Several decisions have not yet been taken. These include the site, where there are three candidates, and the final design, either a triangular shape or two L-shaped interferometers. As for timing, the activities will probably start at the end of the 2030s or the beginning of the 2040s, but this has to be defined more precisely.

What are the challenges of ET? We will have an increase in the event rate and more physical information in the strain signal. Because of the low frequency, we will also have overlapping signals. This will require new algorithms and, of course, more computing power. The low-frequency cutoff will be really important for getting gravitational-wave signals as soon as they are produced, much earlier than is done with second-generation detectors now. We will also need longer templates, which implies a need for more RAM. Virgo's computing is now about 10% of an LHC experiment, and the goal is for ET to stay at roughly the same 10% of an HL-LHC experiment.

To study all of this, the EU funded a preparatory-phase project called ETPP. One of the project's most important results is the delivery of the computing model. The project ends at the end of this year, and it will be followed by a new two-year project called ETCompass. The computing model was recently delivered to the European Commission as one of the deliverables of ETPP. The main question behind it was: can we analyse data that contains 1,000 times more events, including long-lasting events and therefore overlapping signals? The experience of LVK was very important in defining this computing model, as was the expertise of other experiments, including CERN and WLCG.

Most of the LVK computing relies on tools provided by the Open Science Grid, for example HTCondor and Pelican. For some activities, WLCG tools are also used, specifically Rucio for some data transfers. There are already synergies with projects funded under the EOSC umbrella, specifically ESCAPE, but this is only the starting point.

The computing domains are similar to those of other high-energy physics experiments, with one additional domain: rapid alerts. It is really important to detect gravitational waves as soon as possible, in order to send alerts to other telescopes so they can gather additional information at other wavelengths and get more pieces of the picture. The low-frequency cutoff will allow early warning, in many cases much before the merger. The frequency for the Einstein Telescope will start from one or two hertz, so it will be possible to issue these alerts before the merger, by hours or even days. The offline and online domains are more or less similar to what happens in other experiments.

For the alerts, we will deal with 1,000 times more events, some of them overlapping. Currently it takes about 30 seconds to analyse data and generate prompt alerts, using 20,000 cores and roughly 500 GPUs. The increased sensitivity of ET means not only a huge number of additional signals, but also signals that last longer in band. This is a big challenge. We have to improve our methods for detecting signals and eliminate the background more efficiently. For offline, we have to handle this huge number of additional events, which includes improving parameter estimation, testing physics models, testing general relativity, and so on.

For the online domain, the table shows one estimate of the computing needed. In a minimal or safe operational scenario it is not so huge, because what matters most for us is having enough resources for the rapid alerts. If we simply scale what is done today, we estimate a need for 40 million cores for this activity alone. That is of course not possible at all. So we will rely on many different techniques: AI and machine learning, new GPU architectures and other mitigations, to reduce the time and the number of resources we need. You can look at the science blue book, where all of this is described.

For offline, we currently require about 10% of an LHC experiment, and our main goal is to stay at about 10% of an HL-LHC experiment. To do that, we have to use the same techniques as for the rapid alerts.

On data, the most important data are the strain data, time series containing the physics signals. The data volumes are probably similar to what we have in Virgo now, ranging from 10 to 100 terabytes per year. The raw data contain auxiliary channels and sensors and scale with the complexity of the detector, so this will depend strongly on the geometry, triangular or 2L. The expected amount is about 100 petabytes per year. That is significant, but mainly for archives. Only a subset of it is what we have to deal with daily. To first order, most analyses need only the strain data, plus some data-quality information and metadata. All these numbers are overestimated for safety and will be refined in the future.

On sustainability, the computing model includes an estimate of the energy needed. We made assumptions for a minimum and a maximum for low latency, with different PUE values and with CPU-only or mixed GPU and CPU setups. You can see that the lower the PUE, and the more we use a combination of GPUs and CPUs, the better the way forward. We use the same methodology as WLCG for this calculation, but we still have to work out how to make use of all this for sustainable computing. We also have to ensure we have enough resources for an experiment with a lifespan of about 50 years.

The first issue is software. The figure shows how complex the rapid-alert mechanism and the methods we use are. I will not go into the details, but this poses challenges. There is a clear need for a common software framework, best practices for software development, and comprehensive end-to-end testing for software quality. These needs are common, but best practices are especially important for the Einstein Telescope, because several groups are doing different activities, so a common attitude and common practices really matter. We are at the beginning of the story and can rely on past experiments, but there is a lot to do.

We also need intelligent analysis platforms to avoid duplication. We plan to test and reuse existing tools as much as possible, customising them to the needs of the Einstein Telescope. This is already ongoing within the European projects: two funded projects, ETAP and Medden [names unclear], are dealing with these things. Some tools are already being used, such as Rucio for data management and distribution, and the HSF Conditions Database. REANA is used for workflow definitions. There is also a need for monitoring and accounting to predict the needs of the Einstein Telescope, which will include many things of interest for sustainability.

On data centres, in principle ET will not require new ones, but there could be reasons to build some. We already have one project in Germany, which will test a data centre in which computing power adapts to energy availability. The data centre will be housed in a shipping container. It is important to understand to what extent this is a solution that can be used in practice.

Caterina already mentioned the ENSURE project, so I will not go into the details. I think the Einstein Telescope is not officially part of the project, but it is connected through INFN, and through INFN to some Italian national projects. The main role of ET is to provide requirements and needs, test the solutions developed by the project, and provide feedback to the developers. INFN is also involved in this project to connect two national projects funded by the Ministry of University and Research, which provide metrics for new hardware architectures and test solutions. This is quite important, because synergies are always a good approach to addressing very important challenges.

To conclude, the Einstein Telescope is the third-generation gravitational-wave detector. It plans to start operation at the end of the 2030s or the beginning of the 2040s. There are significant computing challenges. The data volumes are quite modest compared with other big experiments, but a 10 to 100 times speed-up in computing is needed, and we plan to stay at about 10% of the capacity of an HL-LHC experiment.

The computing model, which I mentioned earlier, was recently submitted to the European Commission. It is built on the LIGO-Virgo-KAGRA model, but we plan to use WLCG solutions for distributed computing as much as possible.

Sustainability is one of the critical aspects of the ET computing model. There is tight collaboration with WLCG, HSF, KM3NeT, CTA, SKA and Rubin [unclear]. Some of these collaborations we still have to establish, but they are mandatory for us. We will build on the ESCAPE, ETAP and Medden projects I have already mentioned, and on one further project [name unclear]. There is also the green data-centre prototype in Germany. Some funded or proposed projects dealing with environmental sustainability are very important too. In addition to ENSURE, I have to mention one project in preparation called M2 Tech, which will explore how we can address future challenges. Community building and training are also extremely relevant, both in general and for sustainability at large.

That's all from my side for now. Are there any questions?

Markus Schulz: Now I do have a question. You mentioned a lot of collaboration. I just came out of a presentation on LISA, the space-borne gravitational-wave detector. Are you working together with them on potential data processing, or sharing your experience with them?

Paul Laycock: Steven might be in a better position to answer, but there is talk of collaboration at some point. The frequency range is rather different, Markus, so sharing the core analysis of the data streams is not that compelling. At the same time, you can imagine scenarios where interesting objects first appear in LISA and then migrate into the Einstein Telescope. Understanding that handoff, and having a coherent way of tracking the same astronomical object, would be interesting. In terms of data-analysis techniques and algorithms, I personally am not aware of much direct collaboration, but Steven is a lot better connected.

Steven Schramm: We talk about very long signals in ET, lasting hours or days. In LISA, signals can last indefinitely. They will have a background of continuous binaries. Because these are monochromatic and don't change in frequency over decades or more, the longer they analyse the data, the more of them they can see. This makes things a little different. Their approach is more along the lines of a global fit: they iteratively identify and subtract sources, then find more, and so on. For us, the signals are long but not stationary in the same way, even though they last a very long time. So the data analysis is relatively different.

That said, global fits could bring some benefits to ET, because, as was mentioned, there will always be signals present. We will have less ability to dig down into the very weak signals, because there is no same-signal-year-after-year effect where you can just integrate the SNR. It might make sense to combine some kind of hybrid algorithm between what we do now in LVK and what they do in LISA, to get the best of both worlds. But this is very much a future-oriented discussion. I don't think we have any solid examples of it yet, because we don't have the same sensitivity to quasi-stable sources that they do. To answer your question directly, discussions are ongoing. But it's not as simple as it sounds, because the phenomenology of what we are looking for is quite different.

Markus Schulz: If I understood correctly, because you roughly know the frequency window you are looking at, you can use something like sliding-window concepts for the analysis. LISA, on the other hand, has such low frequencies that they basically have to analyse the whole thing continuously in one go.

Steven Schramm: To some extent, yes. We traditionally use matched filters rather than sliding windows, but it is along the lines of what you said. The main difference is that for LISA there are always thousands of signals present. In some parts of the parameter space their background will be unresolved signals rather than noise. It is a physics background that is continuously present. For us, there will always be a mixture of real noise and other signals happening at the same time. As they continue to observe, they can dig deeper and deeper into the physics floor, because they see the same signals recurring over and over. For us, we might have one rotation of the Earth to see something. We will grow the SNR in that time, but we won't reach the level we need for extreme precision.

There have been some studies of hierarchical subtraction, which is basically what LISA does: you identify a signal, subtract it, and re-evaluate what remains in your physics stream. Since we don't have as much precision, each time we do that we introduce errors, and those errors grow pretty quickly. LISA can watch for so long and get such good parameter estimates that there is always some error, but it is not a blocker. For us, it becomes a blocker much more quickly. So we need to find a hybrid between identifying sources, and therefore subtracting the physics background, and matched-filter approaches or their machine-learning equivalents, which can identify individual signals as they occur.

I think this is the biggest unresolved question in data analysis for ET, and we don't know what the outcome will be. It is fairly clear that LISA will need a global fit, and that for LVK matched filtering is the best we have. But in the intermediate regime where ET will work, it's not clear what the solution is. There are definitely discussions going on, but we don't have an answer at this time.

Denisa-Andreea Constantinescu: I have one comment. I wasn't familiar with your work, so I find it very interesting. I was wondering whether there are common types of computation between the SKA Science Data Processor and what you are doing. Any hardware or methods that we design for one telescope could, with some adaptation, be useful for you as well. I don't know exactly what computational algorithms you use to process your data.

Steven Schramm: Luciano, do you want to speak up? I'll say FFTs all the way.

Denisa-Andreea Constantinescu: Okay, then we have something in common.

Steven Schramm: We definitely have some points in common. I joined a bit late, but I think you were talking about ASICs for FFTs, is that correct?

Denisa-Andreea Constantinescu: Yes. ASICs for FFTs have been around for a long time, but we want to build them for non-uniform FFTs. That is something new, and I believe it would be useful. So we are building all the different kernels needed for that, including FFTs.

Steven Schramm: There are definitely potential overlaps there, including with non-uniform FFTs, so we should discuss.

Denisa-Andreea Constantinescu: Okay, let's keep in touch. Do you have any requirements on data format? Are you happy with lower precision than float or double, or something custom?

Steven Schramm: This is a fun discussion. We are discussing it in Virgo right now, because we have once again run into the problem of float32 versus float64. It is partly because of how we store the data. We store the raw value of the strain, which is something times 10⁻²¹. A float32 gets pretty close to the limits of what you can represent, so when you subtract values you go beyond the 32-bit float limit. In the end, I think 32-bit float is probably sufficient for what we need, but we have to be more intelligent about how we store the data. We will probably need some kind of translation layer. Perhaps you store it in 32-bit and convert to 64-bit on the fly. Or you still use 32-bit, but remove the offset, so you are not dealing with something normalised at 10⁻²¹. Done correctly, 32-bit floats are still sufficient for Virgo. I don't know whether that will change for ET. When you get to very small signals we might need something more, and this is also something we are investigating.

Denisa-Andreea Constantinescu: We have been playing a little with the dynamic range of data generated with OSKAR, the virtual SKA telescope simulator. It seems that even 32-bit is too much and you need less, but I don't know what will happen when we get the real data. In hardware, double precision is very costly. I was wondering whether you could go to block floating point, or something even more compressed.

Steven Schramm: I would say it depends on whether you want exact, algorithmic solutions or approximate solutions. Of course, "exact" isn't really a thing when you're talking about floats, but you know what I mean: a descriptive algorithm that you can follow. For that, at least on our side, I think it will be hard to get away from 32-bit. I understand why you would want 16-bit, but it will be challenging for us. Maybe I'm wrong. We need to look into this more, as it's a very recent discussion.

When you go to approximate solutions, such as machine learning, there is a lot of literature showing that you benefit even from float8 and the like. This is not specific to gravitational waves, but generally, because some ambiguity or noise in the system can help it converge better. You might still need the raw data in float32, but all the calculations could be done at a lower floating-point precision, and you could probably still get the vast majority of the performance, if not all of it. There is a lot that remains to be investigated. It is not something we have had to do in the past, but going forward we have to study it more. If you are already looking into this, we could definitely talk.

Denisa-Andreea Constantinescu: Yes, we are studying this, but we don't yet have real data from the telescope to understand the dynamic range at the different stages of the pipeline. We want to see whether we can go lower, and by how much. It also doesn't have to be a standard representation. It can be something between 16 and 32 bits. We need to find the range that is good enough to avoid losing precision, or to keep it at a tolerable level.

Steven Schramm: Fair enough.

Markus Schulz: In my distant past, we developed an ASIC for an application [unclear], and we did not immediately realise that sticking to IEEE was a bad idea. There is also the option of non-linear scales, but then everything you do, if you design a multiplier, gets a bit funny. IEEE is not the end of the universe. You can adapt the number of bits and the amount of exponent space as needed. That is often the biggest advantage of an ASIC.

I think we have already saved a lot of CO2 by bringing the two projects together on developing, and perhaps building and buying, the same ASICs. This is probably more CO2 saved than in any other activity so far. I hope you stay in contact and that something good comes out of this.

Denisa-Andreea Constantinescu: Yes, I would love to.

Disclosure of Delegation to Generative AI (https://panbibliotekar.github.io/gaidet-declaration/)

The authors declare the use of generative AI (GAI) in the research and writing process. According to the GAIDeT taxonomy (2025), the following tasks were delegated to GAI tools under full human supervision:

  • Formatting
  • Summarising

The GAI tool used was: Claude Sonnet 5 (High).

Responsibility for the final manuscript lies entirely with the authors.

GAI tools are not listed as authors and do not bear responsibility for the final outcomes.

Declaration submitted by: Caterina Doglioni

Additional note: started from raw transcript, asked Claude Sonnet 5 (High) to first process it and then summarize in bullet points and action items without adding any information (power consumption of inference only for these two tasks: up to 1 Wh/query according to https://www.sustainabilitybynumbers.com/p/ai-footprint-august-2025, average 2 queries per talk), then in-person pass to correct misunderstandings and remove spurious action items. 

 

There are minutes attached to this event. Show them.