Hands-on reliability work in high-stakes production environments.
Site Reliability Engineer - Ops
PhonePe Pvt. Ltd.
Dec 2025 - Present · Bengaluru, India
Primary on-call responder for mission-critical payment platforms serving 700+ million users.
Built and operated a multi-region active-active MySQL cluster across Azure and two on-prem data centers for high availability and disaster recovery.
Engineered NGINX load-balancing strategies and optimized SQL performance through query profiling and execution plan analysis, reducing p99 latency by 35%.
Centralized database replication monitoring with ELK, Grafana, and Prometheus, reducing root-cause analysis time by 70%.
Monitored YARN ResourceManager and Kubernetes cluster health, optimizing allocation, resolving pending pods, and terminating rogue processes.
Built and maintained Docker and Podman containers with GitLab CI/CD and Harbor for repeatable image builds and production release consistency.
Automated platform health checks, recovery workflows, and incident response with Python and Bash, reducing manual operational effort by 60%.
Built Grafana dashboards on top of Prometheus monitoring to improve SLA/SLO tracking and reduce MTTR by 60% across big data services.
Defined and tracked SLOs and error budgets across owned services, using PagerDuty for alerting and escalation.
Conducted manual chaos and failure-injection testing on staging infrastructure before production rollout.
Product and Platform Engineer (SRE)
Altimetrik Pvt. Ltd. · Client: Visa
Sept 2023 - Dec 2025 · Bengaluru, India (Remote)
Supported 24x7 production operations for Visa's 450 PB Hadoop ecosystem, including HDFS, Hive, and Presto.
Automated AWS infrastructure provisioning and configuration management using Terraform and Ansible, reducing provisioning time by 70%.
Optimized Apache Airflow DAG scheduling and dependency management, improving workflow success rates from 88% to 96% across 500+ daily production pipelines.
Monitored YARN ResourceManager queue-level allocation for Spark jobs and managed Kubernetes cluster health.
Automated platform health checks, recovery workflows, and incident response using Python and Bash, reducing manual operational effort by 60%.
Managed Control-M batch scheduling and CRON jobs, proactively resolving dependency failures and SLA breaches.
Deployed a ServiceNow-integrated L1 support chatbot backed by OpenAI models and historical ServiceNow, Confluence, Jira, and SOP data.
Troubleshot Kerberos authentication, LDAP integration, and IAM policies across the Hadoop ecosystem.
Skills
Tooling across cloud, automation, reliability, and data systems.