Senior Director of Engineering — DevOps & Cloud Platform

Abhishek Anand

14+ years building and scaling cloud infrastructure, Kubernetes platforms, and SRE practices for mission-critical fintech systems — leading 20+ engineers across 10+ business verticals at Paytm, in service of a 300M-user platform.

0+ Years of experience
0K+ Annual cost savings delivered
0+ Services migrated, zero downtime
0+ Professional certifications
Greater Noida, Uttar Pradesh, India · linkedin.com/in/abhishek-anand-k8s
Portrait of Abhishek Anand

About

Platform engineering leadership at fintech scale.

I lead DevOps, Cloud Infrastructure, and Site Reliability Engineering for one of India's largest fintech platforms — translating business and compliance requirements into resilient, cost-efficient infrastructure. My focus is Kubernetes at scale, zero-downtime migrations, security-by-default, and building teams that can operate confidently under national-scale traffic.

  • Leadership — Scaled a DevOps org from 8 to 20+ engineers across 10+ business verticals.
  • Reliability — Improved platform availability from 99.9% to 99.99% at national scale.
  • Cost — Delivered $300K+ in annual infrastructure savings via Graviton, Karpenter & FinOps.
  • Compliance — Led PCI-DSS, SOC 2, and RBI-driven migrations with zero data loss.

Career

Experience

Senior Director of Engineering, DevOps

Apr 2025 — Present

Paytm (One97 Communications Ltd), Noida

  • Introduced agentic AI-powered CI/CD on Amazon EKS, automating canary log and metrics analysis to cut canary-to-live deployment time by 30%.
  • Enhanced EKS MCP server functionality for dev/QA workflows, reducing DevOps dependency for lower environments by 50%.
  • Drove PCI-DSS, SOC 2, and bank partner compliance audits; orchestrated multi-AZ Disaster Recovery drills.
  • Led migration of Paytm domains, repositories, CI/CD pipelines, and EKS manifests to Paytm Payments Services Limited with zero downtime across 5+ major projects.
  • Reduced deployment incidents by 70% using ArgoCD, Argo Rollouts, and Istio service mesh.

Director of Engineering, DevOps

Jul 2024 — Mar 2025

Paytm (One97 Communications Ltd), Noida

  • Architected AWS Graviton migration for 200+ production services, saving $100K+ annually with zero downtime.
  • Achieved 99.99% availability for a 300M+ user platform through Diwali, IPL, and peak-traffic events.
  • Established SRE practices — SLOs, SLIs, error budgets, on-call rotations — across 10+ verticals.
  • Enforced DevSecOps policy-as-code with Kyverno and OPA Gatekeeper across every cluster.
  • Implemented Karpenter-based Spot-to-On-Demand fallback across 5+ EKS clusters with zero interruption.

Senior DevOps Manager

Apr 2023 — Jun 2024

Paytm (One97 Communications Ltd), Noida

  • Executed an emergency 7-day migration of 140+ services from Alibaba Cloud to AWS with zero data loss for RBI compliance.
  • Migrated 100+ microservices from EC2/ECS to Amazon EKS with Helm, RBAC, and namespace isolation.
  • Designed multi-region Disaster Recovery achieving RPO < 5 min and RTO < 15 min.
  • Upgraded EKS, Istio, Kafka, Redis, and Elasticsearch with automated rollback gates and no disruption.

DevOps Manager

Nov 2021 — Mar 2023

Paytm (One97 Communications Ltd), Noida

  • Deployed Karpenter auto-scaling with Kubecost visibility, cutting infrastructure costs 50% ($200K+ annually).
  • Engineered an observability platform (Prometheus, Grafana, Loki, ELK, PagerDuty), lowering MTTR from 45 to 15 minutes.
  • Re-architected CI/CD with Jenkins Shared Libraries and Bitbucket Pipelines, improving release velocity 25%.
  • Attained 100% CIS benchmark compliance, deploying Falco, Kyverno, and Aqua Security across 5+ verticals.

Senior DevOps Engineer & Team Lead

Apr 2018 — Nov 2021

Cloudcover, Pune

  • Delivered 5+ Kubernetes/EC2 migrations from AWS/on-prem to GCP with Helm, ArgoCD, and Spinnaker.
  • Maintained 99.9% uptime SLAs across 10+ client environments on AWS and GCP.
  • Built a multi-tenant vCluster architecture on EKS and AWS Landing Zone, cutting cost and complexity overhead by 70%.
  • Provisioned infrastructure with Terraform, Ansible, and CloudFormation for 50+ services across 8+ projects.

DevOps Engineer

Sep 2015 — Nov 2017

Reliance Jio Money, Bangalore

  • Automated CI/CD with Jenkins for 15+ microservices, shortening release cycles by 80%.
  • Pioneered Docker containerization for dev/QA, eliminating environment drift across 5+ microservices and 20 engineers.

Software Developer

Jun 2012 — Aug 2015

Reliance Jio Infocomm Ltd, Gurgaon

  • Developed the HSS Cluster Manager for Reliance 4G IMS in C++, handling sessions for millions of subscribers.
  • Built foundational expertise in distributed systems and high-availability architecture.

Toolbox

Skills & Technologies

Cloud & Infrastructure

AWSGCPTerraformCloudFormation Disaster RecoveryFinOpsCapacity Planning

Kubernetes & Containers

EKSGKEKOPSDocker HelmIstioMicroservices

CI/CD & GitOps

JenkinsBitbucket PipelinesSpinnaker ArgoCDArgo RolloutsArgo Workflows

Observability

PrometheusGrafanaELK Stack LokiPagerDutySRE / SLOs

Security & DevSecOps

FalcoKyvernoOPATrivy Aqua SecurityCIS Benchmarks

Data, Scripting & Ops

KafkaRedisElasticsearch AnsiblePythonBashLinux

Background

Education

B.Tech (Hons.) in Computer Science

National Institute of Technology (NIT), Jamshedpur

2008 — 2012

Recognition

Accomplishments

  • Paytm Star Performer Award for the 7-day Alibaba-to-AWS migration enabling RBI compliance.
  • CTO Recognition for building the Kubernetes Center of Excellence for the engineering organization.
  • Scaled the DevOps organization from 8 to 20+ engineers through hiring and mentoring.
  • Coordinated the AWS Enterprise Discount Program across 200+ services in 10+ projects.
  • Improved platform reliability from 99.9% to 99.99%, sustaining four-nines availability nationally.

Writing

From the blog

All posts →

Field notes on the real challenges of DevOps and Kubernetes — migrations, cost, reliability, and the trade-offs behind them.

Get in touch

Let's talk platform engineering.

Open to conversations on cloud infrastructure, platform strategy, and engineering leadership.