I'm an AIOps Engineer and Site Reliability Engineer with 5+ years operating and automating production systems across AWS, Azure, and GCP. My focus for the last 3+ years has been AIOps: building observability platforms that don't just show dashboards, but actively correlate signals, cut alert noise, and help teams find and fix problems faster, often before they page anyone.
I work across the full production lifecycle designing Kubernetes and cloud infrastructure, building CI/CD pipelines, running incident response and on-call rotations, and writing the automation that keeps systems self-healing. More recently, that has extended into AI engineering: monitoring LLM/SLM inference workloads, detecting model drift, and building agentic and LLM reasoning based automation for operational tasks.
role: AIOps Engineer / AI Forward Deployed Engineer / SRE
experience: 5+ years in production engineering & platform observability
focus: Observability · Kubernetes · Multi-Cloud · AI-Powered Automation
believes_in: "If it's not observable, it's not production-ready."
|
|
Enterprise client project names are confidential, so the work below is my own engineering built to demonstrate the same production-grade patterns (observability, AIOps, IaC, safe automation) I apply professionally.
|
Production-style multi-cloud observability platform for an agentic AI commerce system. Monitors 4 Kubernetes clusters across AWS, Azure, and GCP, covering 13 AI agents, 7 adversarial "anti-agents," 6 small language models on a GPU serving fleet, and a PCI-scoped payment path entirely defined as Infrastructure as Code.
|
|
An AIOps control loop that reasons about production incidents and fixes them safely. Ingests Prometheus/Alertmanager alerts, classifies them, and decides on a remediation using deterministic playbooks combined with a local LLM (Ollama) reasoning layer then executes it against Kubernetes through a pluggable executor, with every action guarded and audited.
|
|
Serverless order-processing architecture on AWS, built infrastructure-first. A modular serverless system using Lambda, API Gateway, Step Functions, DynamoDB, and SQS to validate, store, and fulfill customer orders with failure handling as a first-class concern, not an afterthought.
|
|
End-to-end CI/CD pipeline: Jenkins, Docker, and AWS EC2.
|
|
Full-stack AI agent product: voice-based, personalized learning experiences.
|
Additional repositories (infrastructure, automation & scripting)
| Repository | Description | Stack |
|---|---|---|
| K8s-Cluster-Setup | Kubernetes cluster bootstrap on a private cloud using kubeadm |
Kubernetes, Shell |
| terraform-aws | Reusable Terraform modules for AWS infrastructure | Terraform, AWS |
| aws-cloudformation-template | CloudFormation templates for EC2, ALB, and VPC | AWS CloudFormation |
| Ansible_playbook_for_aws_instance | Ansible playbooks for configuring AWS EC2 instances | Ansible, AWS |
| infra-setup-php | PHP (Yii2) app deployed via Docker Swarm with an Nginx reverse proxy on EC2 | Docker Swarm, Nginx, AWS |
| cloud_watch | Service for monitoring and managing AWS resources via CloudWatch | AWS CloudWatch, Python |
| Data-extract-from-PDF-file | Extracts structured data from PDFs and loads it into SQL | Python, SQL |
| python-excel-extract | Extracts and aggregates keyword occurrences from Excel comment sections | Python |
- Led an AIOps/SRE pod sustaining 99.95% platform availability across 140+ microservices on multi-cloud Kubernetes, serving 12M+ requests/day
- Cut alert noise by 58% and reduced average MTTR from 45 to 19 minutes by consolidating monitoring into a unified observability stack with AI-assisted anomaly detection
- Redesigned CI/CD pipelines with blue/green and canary gates, taking deployment frequency from twice weekly to daily while cutting rollback incidents by 35%
- Built observability pipelines for AI/ML inference platforms processing 3M+ requests/day across 40+ LLM/SLM-serving services
- Authored Python/Bash self-healing automation removing ~20 engineer-hours/week of manual operational work

