Projects

AI and observability work built around faster incident understanding.

Six projects that balance reliability, automation, and insight — each one shipped or prototyped against a real operational problem. Filter by discipline, or hit ⌘K to jump somewhere else.

LLM-Powered Incident RCA Pipeline

GitHub

An incident RCA pipeline that joins Prometheus metrics with ELK log data and passes the correlated window to an LLM summarisation layer, which drafts runbook recommendations for the on-call engineer.

Impact: contributed to a 35% MTTR reduction during production outages.

Prometheus ELK LLM Incident ops

Transformer-Based Anomaly Detection

GitHub

A TensorFlow Transformer pipeline over Prometheus time-series telemetry that flags latency spikes and infrastructure saturation before they reach customers.

Impact: tuned to 94% precision through iterative threshold calibration.

TensorFlow Prometheus Time series Anomaly detection

On-Prem Incident Response Assistant

Concept

Running Gemma 4 fully offline in a home-lab environment, paired with a log streaming layer that summarises outages and drafts RCA steps without any data leaving the infrastructure.

Impact: a working proof of concept for privacy-first incident response.

Gemma 4 Offline LLM On-prem Automation

Multi-Region MySQL Galera Architecture

Production

An active-active MySQL Galera cluster spanning Azure and two on-premises datacentres, using virtually synchronous replication with explicit failover paths for HA and disaster recovery.

Impact: sub-second synchronisation and zero downtime through node failures.

MySQL Galera Azure HA / DR

Automated Platform Provisioning

Automation

AWS infrastructure provisioning and configuration management automated end to end with Terraform and Ansible, standardising repeatable deployments across platform environments.

Impact: 70% faster provisioning with markedly better deployment consistency.

Terraform Ansible AWS IaC

Airflow Pipeline Reliability Tuning

Data platform

DAG scheduling and dependency management optimised across high-volume production data pipelines, with SLO tracking layered on top of Control-M and Hadoop batch workloads.

Impact: workflow success rate lifted from 88% to 96% across 500+ daily pipelines.

Airflow Control-M Hadoop SLOs

Next

Want the production context behind these?

The experience timeline covers the systems, scale, and on-call reality each project came out of.