Time-Series

1 post in this section

Design a Metrics Monitoring and Alerting System

Monitoring is the system that tells you every other system is broken. Which means when it fails, it fails at exactly the moment you need it — and it fails silently, because the thing that would have told you is the thing that’s down.

We’re building an internal metrics platform: 1,000 server pools, 100 machines per pool, 100 metrics per machine — roughly 10 million metrics, retained for a year.

Continue reading »