This talk will cover recent developments by the CERN security team on their use of Hadoop, Spark and streaming.
Outline:
• Architecture: to explain the goal of the project
• Data ingestion: Flume & Kafka (small intro to why and for what we use it)
• Data processing: The spark jobs we have, problems found, lessons learnt and how do we (want) to monitor them