About

Operational depth, cloud fluency, and a strong bias for automation.

I build and support production systems where downtime is expensive and latency matters. My work spans high-availability infrastructure, incident handling, observability, batch orchestration, secure data platforms, and AI-assisted support flows for teams that need dependable operations.

How I work

I enjoy turning noisy systems into calm ones. In practice that means better alerts, fewer false positives, cleaner runbooks, and automation that helps engineers respond with clarity the moment production signals start to spike.

What I focus on

Reliability engineering, incident response, cloud infrastructure, AI-assisted operations, infrastructure as code, big-data operations, and the unglamorous workflows that quietly reduce toil while improving service health.

Principles

Four rules I keep coming back to.

01

Every alert should be worth waking up for

Alert fatigue is an availability risk. If a page does not map to user impact and a clear first action, it belongs on a dashboard, not in a pager rotation.

02

Automate the second time, document the first

The first occurrence gets a runbook so anyone can resolve it. The second gets a script, a Terraform module, or a self-healing check so nobody has to.

03

Design for the failover you have actually tested

Multi-region only counts once you have injected the failure yourself. Chaos and failure-injection testing in staging is how a DR plan becomes a DR capability.

04

Telemetry is a product, and on-call is the user

Dashboards, log pipelines, and now LLM summarisation exist to answer one question fast: what changed, and who does it affect?

Core strengths

Signal, speed, and steady operations.

24×7 production support SLA & SLO management Incident response Multi-region HA Automation Observability AI for operations Hadoop platform reliability Infrastructure as code Security & access control

Next

See how this shows up in real systems.

Six projects spanning incident RCA, anomaly detection, HA databases, and data pipelines.