Experience

Hands-on reliability work in high-stakes production environments.

On-call leadership, resilient infrastructure, observability, automation, and AI-enabled operations — across payments at PhonePe and a 450 PB data platform for Visa.

Site Reliability Engineer — Ops

PhonePe Pvt. Ltd. · Bengaluru, India

Dec 2025 — Present

  • Primary on-call responder for mission-critical payment platforms serving 700M+ users.
  • Built and operated a multi-region active-active MySQL cluster across Azure and two on-prem datacentres for high availability and disaster recovery.
  • Engineered NGINX load-balancing strategies and tuned SQL through query profiling and execution-plan analysis, cutting p99 latency by 35%.
  • Centralised database replication monitoring across ELK, Grafana, and Prometheus, reducing root-cause analysis time by 70%.
  • Maintained Docker and Podman images with GitLab CI/CD and Harbor for repeatable builds and consistent production releases.
  • Built Grafana dashboards on Prometheus to sharpen SLA/SLO tracking and reduce MTTR by 60% across big-data services.
  • Defined and tracked SLOs and error budgets for owned services, with PagerDuty handling alerting and escalation.
  • Ran manual chaos and failure-injection testing on staging infrastructure ahead of production rollout.

Product & Platform Engineer (SRE)

Altimetrik Pvt. Ltd. · Client: Visa · Remote, Bengaluru

Sept 2023 — Dec 2025

  • Supported 24×7 production operations for Visa's 450 PB Hadoop ecosystem, including HDFS, Hive, and Presto.
  • Automated AWS provisioning and configuration management with Terraform and Ansible, cutting provisioning time by 70%.
  • Optimised Apache Airflow DAG scheduling and dependencies, lifting workflow success from 88% to 96% across 500+ daily pipelines.
  • Monitored YARN ResourceManager queue-level allocation for Spark jobs and managed Kubernetes cluster health.
  • Automated platform health checks, recovery workflows, and incident response in Python and Bash, reducing manual effort by 60%.
  • Managed Control-M batch scheduling and CRON jobs, proactively resolving dependency failures and SLA breaches.
  • Deployed a ServiceNow-integrated L1 support chatbot backed by OpenAI models and historical ServiceNow, Confluence, Jira, and SOP data.
  • Diagnosed and resolved Kerberos authentication, LDAP integration, and IAM policy issues across the Hadoop ecosystem.

Skills

Tooling across cloud, automation, reliability, and data systems.

Bars reflect how often each group shows up in my day-to-day work, not a certification score.

Cloud & infrastructure

88%
  • AWS EC2
  • S3
  • Route 53
  • IAM
  • CloudWatch
  • Azure
  • VPC

Containers & orchestration

92%
  • Kubernetes
  • Docker
  • Podman
  • Helm
  • NGINX
  • Firewalls
  • NAT

Observability & automation

95%
  • Prometheus
  • Grafana
  • ELK Stack
  • Splunk
  • OpenSearch
  • Python
  • Bash

Data & platform

86%
  • MySQL
  • Percona
  • Galera
  • Cosmos DB
  • Aerospike
  • Hadoop
  • Hive
  • Presto

AI & ML

74%
  • TensorFlow
  • LLMs
  • OpenAI models
  • Gemma 4
  • Transformers
  • Anomaly detection

Delivery & operations

90%
  • Jira
  • Confluence
  • ServiceNow
  • Control-M
  • CRON
  • GitLab CI/CD
  • Jenkins
  • PagerDuty

Education & certifications

Formal training plus current cloud certification.

2025

MBA in Data Science

Amity University · 8.4 / 10 CGPA

2022

B.E. Electronics & Communication

CMR Institute of Technology, Bengaluru

Certified

OCI 2025 DevOps Professional

Oracle Corporation · Oracle Cloud Infrastructure

Résumé

Everything above, on one page.

Grab the PDF, or reach out if you would rather just talk through it.