Skip to content
View Jakaria's full-sized avatar

Block or report Jakaria

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Jakaria/README.md

πŸ‘‹ Introduction

Driven by a strong engineering mindset and over a decade of experience in large-scale systems, I specialize in System Architecture, Site Reliability Engineering (SRE), and DevOps. My work focuses on building highly available, scalable, and secure infrastructure that supports mission-critical applications in production environments.

In recent years, I have expanded my expertise into AI Engineering and AI Infrastructure, combining cloud-native principles with modern AI/LLM capabilities. I actively design and build AI-powered systems, leveraging LLMs, Agent Architectures, and Retrieval-Augmented Generation (RAG) to solve real-world business problems.

I thrive at the intersection of Infrastructure + AI, where I can design resilient platforms that not only scale but also intelligently adapt and automate.


πŸš€ Technical Endeavors & Experience

☁️ Cloud Architecture & Infrastructure

  • Design and implement scalable, fault-tolerant, and cost-optimized architectures across AWS, GCP, and Azure
  • Build multi-account, multi-region cloud environments with strong security and governance
  • Hands-on with infrastructure-as-code using Terraform and cloud-native services

πŸ—οΈ System Design & Optimization

  • Architect high-throughput, low-latency systems (15,000+ QPS scale)
  • Identify bottlenecks and implement performance tuning and optimization strategies
  • Design systems with resilience, observability, and scalability as first principles

πŸ“Š Observability & Reliability Engineering

  • Build end-to-end monitoring and alerting systems using tools like Prometheus, Grafana, and Splunk
  • Define and implement SLI/SLO/Error Budget frameworks
  • Drive incident response, root cause analysis, and reliability improvements

πŸ”„ CI/CD & DevOps

  • Design and implement automated CI/CD pipelines for reliable and repeatable deployments
  • Enable GitOps workflows and continuous delivery practices
  • Optimize release processes for speed, safety, and traceability

☸️ Kubernetes & Platform Engineering

  • Deploy and manage Kubernetes (EKS/GKE/AKS) clusters at scale
  • Build internal platforms and developer tooling for improved productivity
  • Optimize container orchestration for performance and cost

πŸ’Ύ Storage & Data Engineering

  • Design storage strategies using S3, RDS, OpenSearch, Data Lakes, and Delta Lake
  • Work with large-scale data ingestion pipelines (e.g., Security Lake, streaming workflows)
  • Optimize storage selection based on performance, durability, and cost requirements

πŸ€– AI Engineering & AI Infrastructure

🧠 LLM & AI System Development

  • Hands-on experience with OpenAI, Claude, Bedrock, and open-source LLMs (Ollama, OpenChat, etc.)
  • Build RAG (Retrieval-Augmented Generation) systems using vector databases (e.g., OpenSearch)
  • Design LLM-powered applications for automation, analytics, and knowledge systems

🧩 Agent Architecture & Agentic Systems

  • Design and implement AI Agents with multi-step reasoning and tool usage
  • Build Agentic workflows integrating APIs, databases, and external systems
  • Experience with frameworks like LangChain / LangGraph and custom orchestration layers

πŸ› οΈ Prompt & Context Engineering

  • Develop advanced prompt engineering strategies for accuracy and reliability
  • Implement context engineering techniques for better grounding and reduced hallucination
  • Optimize LLM outputs for production-grade use cases

⚑ Vibe Coding & AI-Augmented Development

  • Strong advocate of AI-first development workflows
  • Use tools like Claude, Codex, Gemini CLI for planning, design, coding, and debugging
  • Build complete systems (backend + infrastructure) using agentic and AI-assisted workflows

🧱 AI Infrastructure & MLOps Foundations

  • Design infrastructure for LLM deployment, inference, and scaling
  • Integrate AI systems with cloud-native architectures and event-driven pipelines
  • Work with secure, scalable AI pipelines in production environments

🎯 Areas of Interest

  • Cloud Computing (AWS / GCP / Azure)
  • System Design & Distributed Architecture
  • Site Reliability Engineering (SRE)
  • Observability & Performance Optimization
  • CI/CD & DevOps Automation
  • Kubernetes & Platform Engineering
  • AI Infrastructure & LLM Systems
  • Agentic AI & Autonomous Workflows
  • Retrieval-Augmented Generation (RAG)
  • AI-Augmented Development (Vibe Coding)

πŸ’‘ Philosophy

I believe the future of engineering lies in the convergence of Infrastructure and AI.
My goal is to build systems that are not only scalable and reliable, but also intelligent, adaptive, and self-improving.


Popular repositories Loading

  1. python python Public

    Experiment with algorithm, data structures, web applications using Python eco-system

    Python 1

  2. go go Public

    experiment

    Go

  3. Jakaria Jakaria Public

    Config files for my GitHub profile.

  4. nestjs nestjs Public

    Experiment, R&D and Sample Project Using NestJS

    TypeScript

  5. nodejs nodejs Public

    JavaScript

  6. github-action github-action Public

    Experiment and sample pipeline for various kinds of projects