Speaker
Description
Modern software solutions rely on databases that are expected to operate continuously, serve geographically distributed users, and remain trustworthy under constant load. In such environments, a database is no longer a local component or a short‑lived project, but a long‑running service with explicit availability, performance, and reliability expectations.
This lecture explores what fundamentally changes when databases move from development environments to always‑on services. We will examine real‑world constraints such as hardware failure, network instability, data growth, and human error, and how these factors shape operational practices. Topics include monitoring and observability, distinguishing symptoms from root causes, and understanding why many production problems cannot be detected through simple resource metrics alone.
The session emphasizes practical lessons from operating databases at scale, showing why reliability must be designed in from the beginning.