Driven by a strong engineering mindset and over a decade of experience in large-scale systems, I specialize in System Architecture, Site Reliability Engineering (SRE), and DevOps. My work focuses on building highly available, scalable, and secure infrastructure that supports mission-critical applications in production environments.
In recent years, I have expanded my expertise into AI Engineering and AI Infrastructure, combining cloud-native principles with modern AI/LLM capabilities. I actively design and build AI-powered systems, leveraging LLMs, Agent Architectures, and Retrieval-Augmented Generation (RAG) to solve real-world business problems.
I thrive at the intersection of Infrastructure + AI, where I can design resilient platforms that not only scale but also intelligently adapt and automate.
- Design and implement scalable, fault-tolerant, and cost-optimized architectures across AWS, GCP, and Azure
- Build multi-account, multi-region cloud environments with strong security and governance
- Hands-on with infrastructure-as-code using Terraform and cloud-native services
- Architect high-throughput, low-latency systems (15,000+ QPS scale)
- Identify bottlenecks and implement performance tuning and optimization strategies
- Design systems with resilience, observability, and scalability as first principles
- Build end-to-end monitoring and alerting systems using tools like Prometheus, Grafana, and Splunk
- Define and implement SLI/SLO/Error Budget frameworks
- Drive incident response, root cause analysis, and reliability improvements
- Design and implement automated CI/CD pipelines for reliable and repeatable deployments
- Enable GitOps workflows and continuous delivery practices
- Optimize release processes for speed, safety, and traceability
- Deploy and manage Kubernetes (EKS/GKE/AKS) clusters at scale
- Build internal platforms and developer tooling for improved productivity
- Optimize container orchestration for performance and cost
- Design storage strategies using S3, RDS, OpenSearch, Data Lakes, and Delta Lake
- Work with large-scale data ingestion pipelines (e.g., Security Lake, streaming workflows)
- Optimize storage selection based on performance, durability, and cost requirements
- Hands-on experience with OpenAI, Claude, Bedrock, and open-source LLMs (Ollama, OpenChat, etc.)
- Build RAG (Retrieval-Augmented Generation) systems using vector databases (e.g., OpenSearch)
- Design LLM-powered applications for automation, analytics, and knowledge systems
- Design and implement AI Agents with multi-step reasoning and tool usage
- Build Agentic workflows integrating APIs, databases, and external systems
- Experience with frameworks like LangChain / LangGraph and custom orchestration layers
- Develop advanced prompt engineering strategies for accuracy and reliability
- Implement context engineering techniques for better grounding and reduced hallucination
- Optimize LLM outputs for production-grade use cases
- Strong advocate of AI-first development workflows
- Use tools like Claude, Codex, Gemini CLI for planning, design, coding, and debugging
- Build complete systems (backend + infrastructure) using agentic and AI-assisted workflows
- Design infrastructure for LLM deployment, inference, and scaling
- Integrate AI systems with cloud-native architectures and event-driven pipelines
- Work with secure, scalable AI pipelines in production environments
- Cloud Computing (AWS / GCP / Azure)
- System Design & Distributed Architecture
- Site Reliability Engineering (SRE)
- Observability & Performance Optimization
- CI/CD & DevOps Automation
- Kubernetes & Platform Engineering
- AI Infrastructure & LLM Systems
- Agentic AI & Autonomous Workflows
- Retrieval-Augmented Generation (RAG)
- AI-Augmented Development (Vibe Coding)
I believe the future of engineering lies in the convergence of Infrastructure and AI.
My goal is to build systems that are not only scalable and reliable, but also intelligent, adaptive, and self-improving.


