SRE at PhonePe · Bengaluru

Building resilient platforms that stay calm under peak load.

I'm Sudeep Reddy, a Site Reliability & DevOps engineer. I scale distributed platforms, engineer away toil, and use observability plus AIOps to cut MTTR on high-concurrency systems serving 700M+ users.

Years in production
3+
User traffic served
700M+
p99 latency reduced
35%
Kubernetes Terraform Ansible Docker Podman Prometheus Grafana ELK Stack Python Bash Linux MySQL Galera NGINX AWS Azure Airflow Hadoop GitLab CI PagerDuty

What I do

Operations depth, platform breadth, and a bias for clean automation.

01

Platform reliability

Multi-region active-active databases, resilient networking, and production incident workflows across services where downtime is measured in revenue.

02

Cloud & automation

Repeatable infrastructure with Terraform, Ansible, Docker, Podman, Kubernetes, and CI/CD pipelines — faster delivery with safer changes.

03

AI-assisted operations

LLM summarisation over observability data and anomaly-detection workflows that shorten time to insight during live outages.

Impact

Numbers from production, not from a slide deck.

450 PB+ Hadoop ecosystem supported at Altimetrik on the Visa engagement.
70% Faster infrastructure provisioning via Terraform and Ansible.
35% Lower p99 latency through SQL tuning and NGINX optimisation.
96% DAG success rate after tuning 500+ daily Airflow pipelines.

Explore

Focused pages for the parts people care about most.

Résumé

Looking for the complete technical snapshot?

One page, every system, every number — or hit ⌘K to jump anywhere on this site.