Monitoring stack for a fintech platform
The challenge
The client had no centralized visibility into service health — issues were caught by customer complaints, not monitoring.
What I built
A full Prometheus + Grafana stack with custom exporters, SLO-based alerting routed to Slack, and on-call runbooks tied directly to each alert.
Result
Mean time to detection dropped from ~45 minutes to under 5. Incident response time overall fell by 40%.