Monitoring is the system that tells you every other system is broken. Which means when it fails, it fails at exactly the moment you need it — and it fails silently, because the thing that would have told you is the thing that’s down.
We’re building an internal metrics platform: 1,000 server pools, 100 machines per pool, 100 metrics per machine — roughly 10 million metrics, retained for a year.
Continue reading »