Skip to content
Rohit Verma

Open to conversations — platform leadership, SRE & cloud architecture

Senior DevOps Lead · Kubernetes & EKS · AWS · GitOps · SRE · Platform Engineering

I architect and run the platforms that payments ride on — 50+ UPI microservices on AWS carrying 160M+ transactions a day, with the failover, patching and release automation that keeps them standing.

Fifteen years leading DevOps, SRE and platform engineering: Kubernetes and EKS, GitOps with Argo CD and Rollouts, Istio, Terraform, and an observability stack I get paged by — in regulated FinTech, where the failure mode is the front page.

Role
Senior DevOps Lead, Paytm UPI
Based
Noida, India
Experience
15+ years in DevOps, SRE & platform engineering
Fun projects
filehoot.ai · devxops.tech

15+

Years engineering

97.2%

DR failover time cut

160M+

UPI transactions / day

Zero

Downtime on DC → AWS

02

Listen

The same résumé, spoken — for anyone who would rather hear it than read it.

Audio résumé · 3:36

Listen instead of reading

Playback speed
Download audio ↓

Transcript

Introduction

Hello. I'm Rohit Verma, a Senior DevOps Lead based in Noida, India. This is a spoken version of my résumé — about four minutes.

What I do

I architect and run the platforms that payments ride on: 50+ UPI microservices on AWS carrying 160 million transactions a day, with the failover, patching and release automation that keeps them standing. Fifteen years across DevOps, SRE and platform engineering — Kubernetes and EKS, GitOps with Argo CD and Rollouts, Istio, Terraform, and a full observability stack, in regulated FinTech.

Impact

The numbers I'd point to first. Cut DR failover time by 97.2% — architected and automated end-to-end failover runbooks for mission-critical UPI payment services, removing every manual handoff from the recovery path. Engineered fleet-wide patch automation across 200+ EKS nodes and the EC2 estate through custom CI/CD orchestration, delivering zero-downtime patch cycles at scale. Owned the UPI migration from physical DC to AWS Cloud end to end — architecture, networking, data, cutover and post-migration tuning — landing zero downtime and High Availability (HA) across 50+ microservices (160M+ TPD).

Case study — The migration

UPI: physical data centre → AWS, with zero downtime. The entire UPI stack at PPBL (Paytm Payments Bank) had to move from on-premises data centres to AWS. I owned the architecture and the cutover: networking, security, data synchronisation, deployment orchestration and compliance certification (PCI-DSS, ISO 27001, SOC 2) — with no acceptable customer impact at any point.

Case study — Disaster recovery

Full DR movement in under ten minutes. Failover for the UPI stack was a manual runbook — a sequence of human handoffs that took the better part of a shift and was expensive to rehearse. I re-architected the whole path into automation, so a full DR movement now completes end to end in under ten minutes and gets drilled on a schedule rather than survived in an incident.

Case study — Progressive delivery

Releases that roll themselves back. A payment platform can't depend on someone watching a dashboard at 2am to catch a bad release. I designed the delivery path so every deploy goes out as a canary and is judged automatically: Argo Rollouts compares it against the baseline on performance, business and infrastructure metrics, and the release either advances or reverts on its own.

Experience

Currently Senior DevOps Lead at Paytm UPI, since May 2022. Before that: Technical Leader at Capgemini Engineering, Lead Engineer at EVC Ventures, Senior Software Engineer at Appster, Senior Engineer at HCL Technologies, and Software Engineer at Tech Mahindra — fifteen years in all.

Skills

Day to day I work with Kubernetes/EKS, Argo CD, Argo Rollouts (Canary/Experiments), Istio Service Mesh, AWS (EC2, VPC, NLB/ALB, S3, ECR, IAM/IRSA, CloudWatch...), Terraform, Helm, Kafka, Aerospike, Redis, Elasticsearch, Prometheus/Grafana, ELK Stack, Thanos, and the rest of the cloud-native toolchain. I've also earned the Certified Kubernetes Administrator, AWS Solutions Architect Associate, HashiCorp Terraform Associate and Docker Certified Associate certifications.

Get in touch

If you're hiring, migrating, or stuck mid-cutover, write to me at er_rohitverma at yahoo dot in, or find me on LinkedIn. Thanks for listening.

03

Impact

What I changed, and by how much — measured, not adjectival.

200+

EKS nodes patched, zero downtime

50+

Microservices migrated to AWS

30%

P95 latency reduction at peak

25%

Infrastructure cost reduction

60%

Fewer recurring incidents

90%

Provisioning automated via IaC

  1. Cut DR failover time by 97.2% — architected and automated end-to-end failover runbooks for mission-critical UPI payment services, removing every manual handoff from the recovery path.

  2. Engineered fleet-wide patch automation across 200+ EKS nodes and the EC2 estate through custom CI/CD orchestration, delivering zero-downtime patch cycles at scale.

  3. Owned the UPI migration from physical DC to AWS Cloud end to end — architecture, networking, data, cutover and post-migration tuning — landing zero downtime and High Availability (HA) across 50+ microservices (160M+ TPD).

  4. Established progressive delivery on Argo Rollouts (canary/experiments) with KEDA burst autoscaling, driving P95 latency down 30% at peak load.

  5. Standardised automation & IaC on Terraform, Ansible and SaltStack — automating 90% of infrastructure provisioning and cutting incident recurrence 60%.

  6. Hardened the security and compliance posture — IRSA, least-privilege IAM, TLS, CIS baselines, CSPM — while driving a 25% infrastructure cost reduction through FinOps.

  7. Deep hands-on command of Istio, Kafka, Redis, Aerospike, Terraform, Helm, Jenkins, GitHub Actions, ELK Stack, Thanos, OpenTelemetry.

  8. Leads cross-functional engineering — grew and mentored a 5-engineer DevOps team, set the platform roadmap, and steered architecture decisions across 6+ engineering teams.

Architecture & platform initiatives

  • Cut DR failover by 97.2% — fully automated failover runbooks with no manual handoffs; near-instant recovery for UPI payment services.
  • Adopted Istio for service-to-service mTLS, traffic steering and policy-driven security.
  • Standardised CI/CD and GitOps on Argo CD, with safe progressive releases through Argo Rollouts canary and blue-green experiments.
  • Rolled out IRSA and least-privilege IAM patterns, plus CSPM, CIS benchmarks and TLS policy for FinTech compliance.
  • Owned capacity management and KEDA autoscaling for predictable burst handling — P95 latency down 30%.
  • Built SLO/SLA dashboards in Grafana and Thanos, DR runbooks and quarterly failover drills; 25% cost reduction through FinOps.

04

Case studies

Three pieces of work, in full: the constraints, the approach, and what the numbers moved to.

01 · The migration

UPI: physical data centre → AWS, with zero downtime

12 weeks · 50+ microservices · 160M+ transactions/day · Paytm Payments Bank

The entire UPI stack at PPBL (Paytm Payments Bank) had to move from on-premises data centres to AWS. I owned the architecture and the cutover: networking, security, data synchronisation, deployment orchestration and compliance certification (PCI-DSS, ISO 27001, SOC 2) — with no acceptable customer impact at any point.

Technical constraints

  • Entire UPI stack running on legacy on-prem infrastructure.
  • Complex networking: VPC, subnets, peering, NAT/IGW, PrivateLink and DirectConnect for NPCI and bank connectivity.
  • Large-scale data migration with strict consistency requirements.
  • Multi-service dependencies and tight coupling between services.

Business requirements

  • Zero downtime tolerance for UPI transactions.
  • Strict regulatory compliance and audit requirements.
  • Minimal customer impact throughout the migration.
  • Rollback capability at every phase.
  • Obsolete tooling replaced with modern, supported alternatives.
  • Architecture designed for maximum automation, minimum manual intervention.
The plan, phase by phase — and what could have gone wrong
  1. 01 · Weeks 1–3

    Foundation

    • AWS account setup following landing-zone best practices.
    • VPC design: multi-AZ architecture with public/private subnets.
    • Network connectivity: VPN, Direct Connect, peering configurations.
    • Security groups, NACLs, and IAM roles/policies.
    • EKS cluster provisioning with Istio service mesh.
  2. 02 · Weeks 4–6

    Data migration

    • Database replication setup (Aerospike and others).
    • Kafka topic migration and consumer-group synchronisation.
    • S3 buckets with lifecycle policies.
    • Data validation and consistency checks.
    • Performance baselines established.
  3. 03 · Weeks 7–10

    Application deployment

    • Containerisation of all UPI microservices.
    • Helm charts and GitOps setup with Argo CD.
    • Blue/green deployment preparation.
    • Service mesh configuration (Istio routing rules).
    • Observability stack deployment (Prometheus, Grafana).
  4. 04 · Weeks 11–12

    Cutover & validation

    • Progressive traffic shift using Argo Rollouts canaries, gated on automated metric analysis.
    • Real-time monitoring of latency, error rates and throughput.
    • DR runbook execution and failover testing against the automated path.
    • Final cutover during a low-traffic window.
    • Post-migration optimisation and tuning.

Risks & mitigations

Data inconsistency during replication
Dual-write pattern with reconciliation jobs; automated consistency checks before cutover.
Network latency increase
Direct Connect for low-latency connectivity; performance benchmarking at each phase.
Service dependency failure
Circuit breakers, retries with exponential backoff, and comprehensive monitoring dashboards.
Security compliance gaps
IRSA for pod-level IAM, encryption at rest and in transit, audit logging, CIS benchmarks.
Rollback complexity
Automated rollback scripts and fine-grained traffic shifting with Argo Rollouts, rehearsed in drills — with the decision to revert driven by an automated metric comparison rather than a human call mid-incident.

Outcome

100%

Services migrated, zero downtime

-30%

P95 latency

+200%

Autoscaling efficiency

12 wks

End to end

  • Migrated 100% of UPI services with no customer-facing incidents.
  • Established GitOps with Argo CD for declarative infrastructure.
  • Regulatory compliance with full audit trails and security controls.
  • Infrastructure cost reduced through right-sizing and resource efficiency.
  • Comprehensive observability with SLO/SLA dashboards and alerting.

AWS EKS · Kubernetes · Istio · Argo CD · Argo Rollouts · KEDA · Terraform · Helm · Kafka · Aerospike · Prometheus · Grafana · AWS VPC · ALB/NLB · S3 · IAM · Direct Connect · GitOps · Canary · Blue/Green · Service Mesh

02 · Disaster recovery

Full DR movement in under ten minutes

Paytm UPI · Mission-critical payment services · ~6 hrs → < 10 min

Failover for the UPI stack was a manual runbook — a sequence of human handoffs that took the better part of a shift and was expensive to rehearse. I re-architected the whole path into automation, so a full DR movement now completes end to end in under ten minutes and gets drilled on a schedule rather than survived in an incident.

How it works

  • The end-to-end failover runbook is orchestrated as one pipeline: health checks, traffic rerouting, data-sync verification and alerting, in sequence, unattended.
  • Manual handoffs are out of the critical path — recovery no longer waits on the right person being reachable.
  • Data-sync verification gates the movement: the automation proves consistency before it moves traffic, not after.
  • Quarterly failover drills exercise the same automation under real conditions, so the number is rehearsed rather than theoretical.
  • SLO/SLA dashboards in Grafana and Thanos give a live view of recovery state during a drill or a real event.

Outcome

< 10 min

Full DR movement, end to end

97.2%

Reduction in failover time

Zero

Manual handoffs in the path

Quarterly

Drills under real conditions

DR automation · CI/CD orchestration · Kubernetes · EKS · Prometheus · Grafana · Thanos · Alertmanager · Runbook drills

03 · Progressive delivery

Releases that roll themselves back

Paytm UPI · 50+ microservices · Canary + blue/green

A payment platform can't depend on someone watching a dashboard at 2am to catch a bad release. I designed the delivery path so every deploy goes out as a canary and is judged automatically: Argo Rollouts compares it against the baseline on performance, business and infrastructure metrics, and the release either advances or reverts on its own.

How it works

  • Argo CD holds the declarative state; Argo Rollouts drives canary and blue/green steps.
  • AnalysisTemplates query Prometheus at every step, comparing the canary against the baseline on performance, business and infrastructure metrics.
  • Thresholds are fixed up front, so the verdict is deterministic — the same numbers produce the same decision regardless of who is on call.
  • A breach aborts the rollout and reverts traffic automatically; nobody makes a judgement call mid-incident.
  • KEDA and HPA autoscaling absorb the bursts a payment system sees, so the canary is judged under real load rather than a quiet window.

Outcome

Auto

Rollback on metric regression

30%

P95 latency cut at peak

50+

Services on the pipeline

Every

Release gated on metric analysis

Argo CD · Argo Rollouts · AnalysisTemplates · Prometheus · KEDA · HPA · Istio · Canary · Blue/Green · GitOps

05

Experience

Fifteen years, six companies, one throughline: own the platform and make it boring.

  1. May 2022Present

    4 yrs 3 mos

    Senior DevOps Lead

    Paytm UPI · Noida, India

    • Cut DR failover time by 97.2% by architecting end-to-end automation of failover runbooks — eliminating manual handoffs and delivering near-instant recovery for mission-critical UPI payment services.
    • Engineered zero-downtime patch automation across 200+ EKS nodes and the underlying EC2 fleet with custom CI/CD orchestration, eliminating unplanned maintenance windows at scale.
    • Architected and led the zero-downtime UPI migration from on-prem to AWS (50+ microservices, 200+ EKS nodes, 160M+ TPD) — VPC multi-AZ design, DirectConnect/PrivateLink for NPCI and banks, and automated rollback at every step.
    • Championed Istio service mesh (mTLS, traffic steering), standardised CI/CD and GitOps on Argo CD/Rollouts (canary/blue-green), and tuned KEDA/HPA autoscaling to drive 30% lower P95 latency at peak traffic.
    • Drove 90% of infrastructure provisioning into code with Terraform/Ansible/SaltStack; instituted CSPM, audit trails and least-privilege IAM; cut incident recurrence 60% and infrastructure cost 25% through FinOps.
    • Partnered with Security, Compliance and Product to secure PCI-DSS, ISO 27001 and SOC 2 compliance against FinTech regulatory requirements.
    • Built and mentored a 5-engineer DevOps team, defined the platform roadmap against business goals, and shaped architecture decisions across 6+ engineering teams.

    Kubernetes · EKS · Istio · KEDA · Argo CD · Argo Rollouts · AWS · Terraform · Kafka · Aerospike · Prometheus · Grafana · Elasticsearch · Ansible · SaltStack · CSPM · CI/CD · GitOps · DR Automation · Patch Automation · OpenTelemetry

  2. Jun 2020Mar 2022

    1 yr 9 mos

    Technical Leader

    Capgemini Engineering · Gurugram, India

    • Led the container orchestration migration from Docker Swarm to Kubernetes and modernised the CI/CD toolchain.
    • Architected a scalable Python test automation framework executing 30K+ firmware test cases a day.
    • Designed cost-efficient AWS architectures across EC2, Lambda, S3, Route53, ELB, ASG, IAM, CloudWatch, VPC, API Gateway, SQS/SNS and EKS.
    • Automated server configuration and application deployment end to end with Ansible playbooks.

    Kubernetes · Docker · Python · AWS · CI/CD · Ansible

  3. May 2018Sep 2019

    1 yr 4 mos

    Lead Engineer

    EVC Ventures · Gurgaon, India

    • Designed resilient AWS infrastructure for microservices on Docker with Swarm/Kubernetes.
    • Built centralised ELK logging, delivered MERN full-stack features, and established the CI/CD that shortened release cycles.

    AWS · Docker · Kubernetes · ELK · MERN · CI/CD

  4. Oct 2016Apr 2018

    1 yr 6 mos

    Senior Software Engineer

    Appster · Gurgaon, India

    • Set org-level DevOps and test automation strategy, delivering CI/CD pipelines across multiple products.
    • Designed solutions balancing present and future requirements, and built the robust test frameworks behind them.

    DevOps · CI/CD · Test Automation

  5. Mar 2015Oct 2016

    1 yr 7 mos

    Senior Engineer

    HCL Technologies · Noida, India

    • Introduced DevOps practice, built CI/CD pipelines, and led performance monitoring and remediation.
    • Acted as customer-facing SME and primary point of contact for cloud adoption initiatives.

    DevOps · CI/CD · Cloud · Performance Monitoring

  6. Apr 2011Mar 2015

    3 yrs 11 mos

    Software Engineer

    Tech Mahindra · Noida, India

    • Built and deployed UI and API test automation frameworks on open-source tooling.
    • Introduced CI/CD and version-control best practice; maintained applications on VMware ESXi.

    Test Automation · CI/CD · VMware ESXi

06

Skills

Run in production, under real traffic, at payment-grade stakes.

Orchestration & Cloud
Kubernetes/EKSAWS (EC2, VPC, NLB/ALB, S3, ECR, IAM/IRSA, CloudWatch...)HelmKEDA AutoscalingMicroservices
Delivery & GitOps
Argo CDArgo Rollouts (Canary/Experiments)TerraformGitOpsJenkinsAnsibleSaltStackInfrastructure as Code (IaC)
Networking & Mesh
Istio Service MeshNginxHAProxyHigh Availability (HA)
Data & Streaming
KafkaAerospikeRedisElasticsearch
Observability
Prometheus/GrafanaELK StackThanosAlertmanager
Reliability & Cost
SRE/ResilienceDR/FailoverCost OptimizationData-Driven Decisions
Languages & Automation
PythonGolang (K8s controllers)Automation

07

Fun projects

Built for the fun of it — shipped, live, and maintained past launch day.

  • Continuously updated DevXOps news platform—stay current with the latest in DevOps, Cloud, Security, and AI operations.

    • Real-time updates on DevOps trends, tools, and best practices.
    • Organized by categories: DevCloudOps, DevSecOps, DevAIOps, and more.
    • Snackable content cards with sources, images, and search functionality.
    • Hands-on project for mastering Next.js, content management, and modern web patterns.

    DevOps · Next.js · News/Content · Real-time

  • Browser-based tool for working with structured data files—converts, validates, repairs, and visualizes JSON, YAML, TOML, and CSV entirely in-browser with no data leaving your device.

    • Format conversion: bidirectional conversion between JSON, YAML, TOML, CSV, and plain text.
    • AI-powered repair: fixes broken JSON/YAML/TOML using browser-based LLMs.
    • Validation with syntax checking, error detection, and line/column reporting.
    • Visualization: Tree, Table, and Raw views with syntax highlighting.
    • AI assistant (Hoot): answers questions about your data using WebLLM.
    • Full-text search with match highlighting—all processing happens locally for privacy.

    AI · WebLLM · Data Tools · Browser-based · Privacy-first

08

Credentials

Certified, and kept current.

Education

  • B.Tech., Electrical, Electronics & Communications Engineering

    Dr. A.P.J. Abdul Kalam Technical University

    20062010 · First

  • Higher Secondary

    Boys' High School & College

Certifications

  • CKA: Certified Kubernetes Administrator

    The Linux Foundation

    2022-03 · Earned (valid to 2025-03)

  • AWS Certified Solutions Architect – Associate

    Amazon Web Services

    2022-01 · Earned

  • HashiCorp Certified: Terraform Associate

    HashiCorp

    2022-03 · Earned

  • Docker Certified Associate

    Mirantis

    2022-02 · Earned

09

Contact

Hiring, migrating, or mid-cutover at 3am — write to me.

Protected by reCAPTCHA — Google Privacy Policy and Terms apply

Location · Response
Noida, India · within 24–48 hours

Take the résumé with you

Résumé ↓Write to me