Ajay Akilan
Site Reliability Engineer — DevOps, platform engineering
Summary
Site reliability engineer with five years operating Zoom’s global instant-messaging platform — billions of requests daily on multi-cloud infrastructure spanning AWS, Azure and OCI. Specialised in migrating live production systems without downtime, building reliability and release tooling in Python, and delivering multi-region active-active architectures. Software development combined with full production ownership: on call, incident command, capacity and cost.
15 production clusters · CKA certified · the platform has sustained 99.99% availability · in five years no customer-impacting incident has been caused by a change of mine, with a 100% rollback success rate
Experience
Zoom Video Communications
July 2021 — presentSite Reliability Engineer (DevOps), Instant Messaging Platform — Bangalore, India
Service owner for Zoom’s core instant-messaging microservices — backend APIs, real-time messaging, and the public edge — with end-to-end production responsibility for releases, incidents, capacity and cost. Scope grew from single-service operations to running fleet-wide infrastructure programmes.
Zero-downtime infrastructure migration
- Led the replatform of the public-facing edge layer across the global production fleet on two clouds, using phased traffic cutovers with automated verification and failover validation at each stage — no customer impact, and edge infrastructure cost on the largest cluster cut in half.
- Owned the move of the messaging microservices from virtual machines onto Kubernetes within a company-wide containerisation programme — built the infrastructure definitions and deployment pipelines with verification embedded, cutting deploy time from hours to under 3 minutes and rollback to under 2.
- Ran the cache-layer transition from Redis to Valkey as a documented, repeatable process adopted as the team standard, and retired the legacy edge stack through staged decommission with full DNS, certificate and monitoring cleanup.
Reliability engineering through software
- Built the automated release-verification framework for the messaging services — from shell and Ansible scripts to gates native to the deployment platform, validating every pod on every production release — and co-developed the parallelised deployment workflow that together cut the multi-cluster release cycle from ~7 hours to 52 minutes.
- Wrote the Python CLI for production DNS operations on a 100,000+ record zone: resolution-chain tracing with live health-check overlay, bulk weight changes behind a preview–confirm–apply flow, and per-change audit logging. Primary tooling for all traffic cutovers; packaged and distributed to DevOps teams across the company.
- Automated the TLS certificate lifecycle for the global edge — zero-downtime rotation validated on isolated pre-production nodes — and built Python utilities for multi-repository management, cloud automation and log analysis, removing an estimated 10–15 hours of manual work weekly.
Multi-region architecture and multi-cloud operations
- Ran the active-active rollout for the chat backend and edge services across 14 of 15 production clusters — dual-region live traffic with weighted DNS routing and failover integrated with an internal traffic-orchestration console — at near-zero recovery time.
- Built production clusters from the ground up on three clouds, including one region delivered end to end: infrastructure, five services, active-active enablement and edge migration. Earlier builds on segregated accounts met government data-residency and compliance requirements.
- Incident commander for the messaging estate, on call from the third month of tenure: hundreds of P0 and P1 events resolved alongside NOC, development and platform teams, each turned into monitoring, alerting or runbook improvement.
Additional scope
- Drove cost optimisation across owned services through instance right-sizing, storage-mode tuning and Kubernetes resource refinement — six-figure annual savings with per-cluster reductions of 29–77%; authored the ongoing optimisation roadmap.
- Took the messaging services through the company's observability platform migration — hundreds of dashboards, severity-tiered alerting tied to runbooks, and SLI/SLO definitions set with the development team.
- Took two newly launched services through to general availability as lead DevOps owner; mentored 3–4 engineers and served as the standing escalation point across India, US and China engineering teams.
- Executed ~90% of global production releases from India through staged waves, establishing the India team's full operational independence.
Skills
- Cloud
- AWS — EKS, EC2, Route 53, DynamoDB, ElastiCache, VPC, IAM, Lambda, S3, CloudFormation · Azure — AKS, Cosmos DB, Functions, Azure Cache for Redis · Oracle Cloud Infrastructure
- Containers and orchestration
- Kubernetes (CKA), Docker · requests and limits, resource quotas, autoscaling, namespace conventions, RBAC, service topology, ingress
- CI/CD and delivery
- Jenkins, GitLab CI, Tekton, an internal Kubernetes delivery platform · canary, blue-green, rolling, active-active, isolated pre-production nodes · automated verification gates, parallel deploy and rollback
- Infrastructure as code
- Terraform (remote state, plan-review-apply), Ansible · protected branches, mandatory peer review, ticket-linked changes
- Observability
- Prometheus, Grafana, Elasticsearch and Kibana, an internal metrics platform built on Datadog · SLI/SLO definition, severity-tiered alerting tied to runbooks, log-based alerting, process monitoring, incident command
- Networking and edge
- NGINX, OpenResty and Lua, NGINX Plus · NLB/ALB, DNS at scale with weighted and health-checked routing, VPC peering, TLS certificate lifecycle
- Data and caching
- DynamoDB, Cosmos DB, Redis, Valkey, ElastiCache · capacity-mode and storage-tier tuning, cache sizing
- Languages
- Python (boto3, CLI tooling, automation), Bash, Lua · Linux production debugging (strace, netstat, lsof)
Education
Certifications
Languages
English — full professional. Tamil — native.