112th ROOT Parallelism, Performance and Programming Model Meeting

Europe/Zurich
Enric Tejedor Saavedra (CERN), Enrico Guiraud (EP-SFT, CERN)
Description
    • 16:00 17:00
      Evaluating Query Languages and Systems for High-Energy Physics Data 1h

      PAPER
      https://arxiv.org/abs/2104.12615

      AUTHORS
      Dan Graur (1), Ingo Müller (1), Mason Proffitt (2), Ghislain Fourny (1), Gordon T. Watts (2), Gustavo Alonso (1) ((1) Department of Computer Science, ETH Zurich, (2) Department of Physics, University of Washington)

      ABSTRACT
      In the domain of high-energy physics (HEP), query languages in general and SQL in particular have found limited acceptance. This is surprising since HEP data analysis matches the SQL model well: the data is fully structured and queried using mostly standard operators. To gain insights on why this is the case, we perform a comprehensive analysis of six diverse, general-purpose data processing platforms using an HEP benchmark. The result of the evaluation is an interesting and rather complex picture of existing solutions: Their query languages vary greatly in how natural and concise HEP query patterns can be expressed. Furthermore, most of them are also between one and two orders of magnitude slower than the domain-specific system used by particle physicists today. These observations suggest that, while database systems and their query languages are in principle viable tools for HEP, significant work remains to make them relevant to HEP researchers.

      Speaker: Ingo Müller (ETH Zurich)