Platform reliability
Multi-region active-active databases, resilient networking, and production incident workflows across services where downtime is measured in revenue.
SRE at PhonePe · Bengaluru
I'm Sudeep Reddy, a Site Reliability & DevOps engineer. I scale distributed platforms, engineer away toil, and use observability plus AIOps to cut MTTR on high-concurrency systems serving 700M+ users.
What I do
Multi-region active-active databases, resilient networking, and production incident workflows across services where downtime is measured in revenue.
Repeatable infrastructure with Terraform, Ansible, Docker, Podman, Kubernetes, and CI/CD pipelines — faster delivery with safer changes.
LLM summarisation over observability data and anomaly-detection workflows that shorten time to insight during live outages.
Impact
Explore
Background, reliability mindset, and the engineering principles behind the work.
Read moreAI incident response, anomaly detection, data-platform reliability, and observability.
See the workPhonePe and Visa work, skills, education, and certifications in one place.
View timelineRésumé
One page, every system, every number — or hit ⌘K to jump anywhere on this site.