- Primary on-call responder for mission-critical payment platforms serving 700M+ users.
- Built and operated a multi-region active-active MySQL cluster across Azure and two on-prem datacentres for high availability and disaster recovery.
- Engineered NGINX load-balancing strategies and tuned SQL through query profiling and execution-plan analysis, cutting p99 latency by 35%.
- Centralised database replication monitoring across ELK, Grafana, and Prometheus, reducing root-cause analysis time by 70%.
- Maintained Docker and Podman images with GitLab CI/CD and Harbor for repeatable builds and consistent production releases.
- Built Grafana dashboards on Prometheus to sharpen SLA/SLO tracking and reduce MTTR by 60% across big-data services.
- Defined and tracked SLOs and error budgets for owned services, with PagerDuty handling alerting and escalation.
- Ran manual chaos and failure-injection testing on staging infrastructure ahead of production rollout.
Experience
Hands-on reliability work in high-stakes production environments.
On-call leadership, resilient infrastructure, observability, automation, and AI-enabled operations — across payments at PhonePe and a 450 PB data platform for Visa.
- Supported 24×7 production operations for Visa's 450 PB Hadoop ecosystem, including HDFS, Hive, and Presto.
- Automated AWS provisioning and configuration management with Terraform and Ansible, cutting provisioning time by 70%.
- Optimised Apache Airflow DAG scheduling and dependencies, lifting workflow success from 88% to 96% across 500+ daily pipelines.
- Monitored YARN ResourceManager queue-level allocation for Spark jobs and managed Kubernetes cluster health.
- Automated platform health checks, recovery workflows, and incident response in Python and Bash, reducing manual effort by 60%.
- Managed Control-M batch scheduling and CRON jobs, proactively resolving dependency failures and SLA breaches.
- Deployed a ServiceNow-integrated L1 support chatbot backed by OpenAI models and historical ServiceNow, Confluence, Jira, and SOP data.
- Diagnosed and resolved Kerberos authentication, LDAP integration, and IAM policy issues across the Hadoop ecosystem.
Skills
Tooling across cloud, automation, reliability, and data systems.
Bars reflect how often each group shows up in my day-to-day work, not a certification score.
Cloud & infrastructure
88%- AWS EC2
- S3
- Route 53
- IAM
- CloudWatch
- Azure
- VPC
Containers & orchestration
92%- Kubernetes
- Docker
- Podman
- Helm
- NGINX
- Firewalls
- NAT
Observability & automation
95%- Prometheus
- Grafana
- ELK Stack
- Splunk
- OpenSearch
- Python
- Bash
Data & platform
86%- MySQL
- Percona
- Galera
- Cosmos DB
- Aerospike
- Hadoop
- Hive
- Presto
AI & ML
74%- TensorFlow
- LLMs
- OpenAI models
- Gemma 4
- Transformers
- Anomaly detection
Delivery & operations
90%- Jira
- Confluence
- ServiceNow
- Control-M
- CRON
- GitLab CI/CD
- Jenkins
- PagerDuty
Education & certifications
Formal training plus current cloud certification.
MBA in Data Science
Amity University · 8.4 / 10 CGPA
B.E. Electronics & Communication
CMR Institute of Technology, Bengaluru
OCI 2025 DevOps Professional
Oracle Corporation · Oracle Cloud Infrastructure
Résumé
Everything above, on one page.
Grab the PDF, or reach out if you would rather just talk through it.