How I work
I enjoy turning noisy systems into calm ones. In practice that means better alerts, fewer false positives, cleaner runbooks, and automation that helps engineers respond with clarity the moment production signals start to spike.
About
I build and support production systems where downtime is expensive and latency matters. My work spans high-availability infrastructure, incident handling, observability, batch orchestration, secure data platforms, and AI-assisted support flows for teams that need dependable operations.
I enjoy turning noisy systems into calm ones. In practice that means better alerts, fewer false positives, cleaner runbooks, and automation that helps engineers respond with clarity the moment production signals start to spike.
Reliability engineering, incident response, cloud infrastructure, AI-assisted operations, infrastructure as code, big-data operations, and the unglamorous workflows that quietly reduce toil while improving service health.
Principles
Alert fatigue is an availability risk. If a page does not map to user impact and a clear first action, it belongs on a dashboard, not in a pager rotation.
The first occurrence gets a runbook so anyone can resolve it. The second gets a script, a Terraform module, or a self-healing check so nobody has to.
Multi-region only counts once you have injected the failure yourself. Chaos and failure-injection testing in staging is how a DR plan becomes a DR capability.
Dashboards, log pipelines, and now LLM summarisation exist to answer one question fast: what changed, and who does it affect?
Core strengths
Next
Six projects spanning incident RCA, anomaly detection, HA databases, and data pipelines.