Vijay Mourya
Platform Engineer — Developer Platforms & AI Tooling

Professional Experience

Seven years building internal developer platforms, multi-cloud infrastructure, and production-grade CI/CD systems across AWS and GCP.

Career Highlights

7
Years of Experience
3
Major Organizations
1,000+
Weekly CI Jobs
500+
AWS Accounts Served

Key Achievements

🎤 Roche representative at AWS re:Invent 2025

Represented Roche at AWS re:Invent 2025, collaborating with AWS TAMs on service roadmaps.

🏆 Technical Excellence

Multiple internal recognitions for AI tooling innovation and multi-cloud platform leadership.

🤖 AI-Assisted Development

Built an autonomous code-review agent and a Bedrock-backed log analysis CLI used by the engineering team.

📚 Documentation & Enablement

Technical documentation, SOPs, video tutorials, and a documentation hub for global engineering teams.

Conference Attendance

☁️
AWS re:Invent 2025
Las Vegas, USA
🌐
KubeCon + CloudNativeCon
India 2025 • Hyderabad

Career Timeline

Detailed work experience in chronological order (most recent first)

Roche Information Solutions India
Pune, India
Jul 2023 – Present
Present
DevOps Engineer — [Acting] Lead Platform Engineer & Service Owner
Technical Leadership & Service Ownership
  • Acting Team Lead for a globally distributed team of 7 engineers, taken on during a team transformation
  • Helped build the team from the ground up, expanding architectural scope from AWS to multi-cloud
  • Own product roadmap, delivery timelines, and stakeholder management, working closely with the Product Manager to deliver a unified multi-cloud platform to three enterprise stakeholder groups
  • Manage technical hiring and onboarding, including international onboarding conducted in Switzerland
  • Mentor engineers and resolve technical blockers day to day
AI Tooling & Developer Productivity
  • Developed an autonomous code-review agent using ONA that analyzes pull requests, suggests codebase refactoring, and asks reviewers targeted questions to improve code quality
  • Built an AI-powered CLI with Amazon Bedrock for log analysis, reducing hours of debugging to minutes
  • Wrote custom Python data pipelines that clean, deduplicate, and trim log streams before the LLM call, keeping token costs low while improving accuracy and speed
  • Developed a Proof of Concept RAG pipeline on AWS to enhance the internal documentation portal
Zero-Trust Self-Service IAM & Kubernetes CI/CD Platform
  • Replaced slow IAM access request tickets with a self-service, Terraform-based control plane where developers obtain temporary or permanent AWS access via OIDC, governed entirely by GitOps MR approvals instead of static keys
  • Architected and developed a centralized, autoscaling GitLab Runner platform on AWS EKS (Fargate + Karpenter) and GCP GKE
  • Delivered tenant-isolated build environments that securely process 1,000+ weekly CI jobs across 5 engineering teams
  • Leading the migration of CI pipelines and codebase from GitLab to GitHub
Centralized Multi-Cloud OS Image Factory & Compute Operations
  • Co-architected a multi-cloud OS image factory serving 500+ AWS and 100+ GCP accounts, owning the full image lifecycle: build, patch, release, communicate, support
  • Built a fully automated, dynamically generated build, test, and hardening pipeline with automated release notes and stakeholder email communication
  • Engineered pipelines provisioning 5 OS families and 20+ compute variants tailored for Base, EKS, ECS, GKE, Deep Learning, and Parallel Cluster workloads
  • Set up specialized build pipelines for HPC/ParallelCluster and GPU-based Deep Learning images, modifying CIS hardening rules to support Slurm scheduling and complex ML networking
  • Built a serverless AMI usage tracker and automated cleanup process retiring unused images, giving the team compliance visibility and saving roughly $10,000 in cloud storage costs
  • Built self-service pipelines for cross-org image copying and SSM document-based server management
  • POC self-service image-building CLI enabling anyone at Roche to customize an OS image while staying within compliant security standards
Serverless Patching Service for AWS
  • Architected an event-driven, set-and-forget serverless patching platform (Lambda, EventBridge, SQS) automating monthly updates for 1,000+ EC2 instances across 500+ AWS accounts, reducing one week of manual work to one day
  • Designed advanced governance controls and guardrails with pre/post-patch hooks and patch promotions, automatically filtering out ASG/EKS clusters to prevent disruptive reboots
  • Took the service from MVP to adoption by 10+ product teams
  • Developed a statistics and data processing pipeline tracking adoption metrics, using those insights to drive new feature development based on real user pain points
  • Managed the AWS account management pipeline, implementing Service Control Policies (SCPs) and IAM hardening based on security protocols
Advocacy, Enablement & Communication
  • Represented Roche at AWS re:Invent 2025 (Las Vegas) and KubeCon 2025 (Hyderabad), working directly with AWS Technical Account Managers (TAMs) to align internal architecture with new cloud features
  • Conduct technical presentations for major platform releases and architecture updates to 200+ technical and business stakeholders
  • Building a documentation hub for the team's services to strengthen team representation and the support model
  • Directed monthly release cycles and recorded technical tutorial videos to streamline team and end-user onboarding
  • Author technical and non-technical documentation and communications for cross-functional alignment
Tech Stack: AWS (EKS, Fargate, Karpenter, Lambda, EventBridge, SQS, Systems Manager, Bedrock, Athena, S3) • GCP (GKE, Cloud Build) • ONA • Claude Code • GitHub Copilot • MCP • Kubernetes • KEDA • Terraform • Terragrunt • OIDC • GitLab CI/CD • GitHub Actions • Packer • Ansible • InSpec • Helm • Docker • Python • Jinja2 • Bash
Amazon Development Centre India
Hyderabad, India
Mar 2022 – Jun 2023
1 year 3 months
DevOps engineer, On-Call PoC
Professional Experience
  • High-Scale Data Engineering: Designed and implemented high-throughput AWS Lambda workflows to process TB-scale CSV/JSON data from S3
  • Lambda Optimization: Overcame execution timeout constraints by leveraging asynchronous invocation, file indexing, and function chaining to optimize data ingestion into DynamoDB
  • Infrastructure Migration: Orchestrated a large-scale HTTP(s) VIP migration, transitioning legacy infrastructure from NetScaler to AWS native load balancers (ALB/NLB)
  • Migration Strategy: Developed comprehensive rollback and automation strategies, including IAM role design, prerequisite validation, and migration timeline planning
  • Incident Management: Served as the Primary On-Call POC for high-severity incidents, executing mitigation strategies and Root Cause Analysis (RCA) to minimize customer impact
  • Operational Excellence: Authored detailed technical documentation, SOP guide books, and mitigation playbooks using Draw.io, which expedited team onboarding and reduced recurring incidents
  • Observability & Monitoring: Led KPI-driven monitoring initiatives, identifying critical metrics and anomalies to enhance proactive system alerting
  • Serverless Automation: Engineered AWS-based serverless tools for organizational ticket tracking and process automation, improving overall operational scalability
  • Global Collaboration: Partnered with globally distributed teams to conduct system design reviews and ensure production-ready code quality for robust deliverables
Tech Stack: AWS (Lambda, S3, DynamoDB, ALB/NLB, CloudWatch, SNS, SQS) • Python • Draw.io • On-Call Operations • Root Cause Analysis
Tata Consultancy Services
Nagpur, India
Jul 2019 – Mar 2022
2 years 9 months
DevOps Engineer
Professional Experience
  • Cloud Migration & IaC: Provisioned scalable AWS compute and storage resources using Terraform and CloudFormation to support large-scale cloud migration projects
  • Infrastructure Automation: Automated custom AMI creation and tool-baking processes using Packer and Ansible, ensuring environment consistency across the organization
  • Disaster Recovery (DR) Engineering: Engineered "point-in-time" disaster recovery solutions; developed automated failover mechanisms using Jenkins jobs triggered by SQS and Lambda-based monitoring
  • CI/CD Orchestration: Managed robust Jenkins CI/CD pipelines for microservices and Micro Frontends (MFEs), leveraging shared libraries and multibranch pipelines integrated with AWS CLI and SSM
  • Observability & Alerting: Automated infrastructure monitoring using CloudWatch, SNS, and Lambda to ensure high availability and rapid incident response
  • Client Management & SOPs: Facilitated infrastructure knowledge transfers and project updates for clients; authored documentation and SOPs to standardize Change Management and onboarding processes
Tech Stack: AWS (EC2, S3, CloudWatch, SNS, SQS, Lambda) • Terraform • CloudFormation • Jenkins • Packer • Ansible • Docker • Kubernetes • Python • Bash

Complete Technical Skillset

🤖 AI & Developer Tooling
  • Autonomous code-review agents (ONA)
  • LLM Integration (Amazon Bedrock)
  • Claude Code, GitHub Copilot, MCP servers
  • RAG Pipeline implementation
  • Prompt & token-cost engineering for DevOps logs
☁️ Cloud Platforms
  • AWS (Advanced): EKS, Fargate, Karpenter, Lambda, EventBridge, SQS, Systems Manager, Bedrock, S3, DynamoDB, Athena
  • GCP (Intermediate): GKE, Compute Engine, Cloud Functions, Cloud Storage
  • Multi-cloud architecture across 500+ AWS and 100+ GCP accounts
  • Cloud cost optimization
🏗️ Infrastructure as Code & Imaging
  • Terraform (modules, remote state, workspaces)
  • Terragrunt
  • Ansible (playbooks, roles, vault)
  • Packer (hardened OS image factories, CIS hardening)
  • InSpec (compliance testing)
  • Helm (chart development & management)
  • CloudFormation
🚀 CI/CD & Platform Engineering
  • GitLab CI/CD (KEDA, tenant-isolated autoscaling runners)
  • GitHub Actions & GitLab → GitHub migration
  • ArgoCD / GitOps practices
  • Security scanning (Trivy, SonarQube)
  • Self-service developer platforms
🐳 Containers & Orchestration
  • Kubernetes / CKA (EKS, GKE, strict tenant isolation)
  • EKS on Fargate with Karpenter autoscaling
  • KEDA (event-driven autoscaling)
  • Docker (multi-stage optimization)
  • HPC / ParallelCluster (Slurm) and GPU Deep Learning images
  • Container security & best practices
🔒 Security & Governance
  • Zero-trust IAM & OIDC federation
  • Self-service access control planes
  • Service Control Policies (SCPs) & IAM hardening
  • CIS hardening & compliance automation
  • Patch governance with pre/post-patch hooks
📊 Observability & Reliability
  • Prometheus & Grafana
  • CloudWatch (metrics, logs, alarms)
  • KPI-driven alerting & Root Cause Analysis
  • SRE practices & SLA/SLO management
📚 Documentation & Communication
  • Technical and non-technical documentation
  • Documentation portals & support models
  • Technical presentations to 200+ stakeholders
  • Tutorial video production (OBS Studio)
  • Architecture diagrams (Lucidchart, Draw.io)