Skip to content
View baratamavinash225's full-sized avatar
  • Lowes
  • Bengaluru

Block or report baratamavinash225

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
baratamavinash225/README.md

Hey there πŸ‘‹, I'm Avinash

Senior Data Engineer | Backend Engineer | GCP | Big Data | ML & GenAI


πŸš€ About Me

  • πŸ’Ό 11.5+ years of experience in Data Engineering, Backend Development & Machine Learning
  • ☁️ Strong focus on Cloud (GCP, AWS, Azure) and scalable data platforms
  • πŸ”„ Expertise in building end-to-end data pipelines (batch & streaming)
  • 🧠 Experience in ML pipelines, NLP, and data-driven systems
  • πŸ—οΈ Passionate about designing production-grade, scalable architectures
  • ⚑ Focus on performance, cost optimization, and data reliability

🏒 Current Focus

  • Building GCP-based data platforms using BigQuery, Dataproc & Composer
  • Developing data observability & auditing systems (usage tracking, cost insights)
  • Designing backend APIs & microservices for data systems
  • Working on streaming + batch architectures (Lambda/Kappa)

☁️ Cloud & Data Platforms

Google Cloud (Primary)

  • BigQuery, GCS, Dataproc, Cloud Composer (Airflow), Pub/Sub
  • Vertex AI (ML workflows & pipelines)

AWS

  • S3, SNS, SQS, EventBridge, EMR, EKS

Azure

  • ADLS, ADF, Databricks, Synapse

βš™οΈ Data Engineering Stack

  • Big Data: Spark, Hadoop, Hive, Kafka, Kafka Connect, NiFi
  • Processing: Batch & Streaming (Lambda Architecture)
  • Orchestration: Airflow, Cloud Composer, Oozie, Dolphin Scheduler
  • Data Stores: BigQuery, PostgreSQL, MySQL, Oracle, HBase
  • Pipelines: ETL / ELT, Data Warehousing, Data Lakes

πŸ–₯️ Programming & Backend

  • Backend: Spring Boot, Microservices, REST APIs
  • Languages: Python, Java, Scala
  • System Design: Scalable distributed systems

🧠 ML, NLP & GenAI

  • Feature Engineering & Data Preparation
  • NLP-based systems (Entity Resolution, Text Processing)
  • Supervised & Unsupervised Learning
  • Deep Learning & Neural Networks
  • MLOps & pipeline integration
  • Exploring LLMs, embeddings, and GenAI workflows

πŸ› οΈ DevOps & Infrastructure

  • Docker, Kubernetes
  • CI/CD (Jenkins, GitOps)
  • Ansible, Rundeck
  • Linux, Shell scripting
  • Secure systems with Vault

πŸ“Š Key Strengths

  • βœ… Designing scalable data platforms
  • βœ… Building reliable data pipelines (TB-scale)
  • βœ… Optimizing BigQuery cost & performance
  • βœ… Developing backend systems for data applications
  • βœ… Enabling data governance & observability

πŸ“Œ Highlight Projects

  • πŸ” Secure Secret Management (Vault + Kafka Connect)
  • πŸ“Š BigQuery Usage Audit Platform (Data Observability & FinOps)
  • 🧠 Entity Resolution using NLP for Corporate Clients

πŸ“« Let's Connect


⭐️ Always open to collaborating on Data Engineering, Cloud, and AI-driven projects

Popular repositories Loading

  1. Seek_Data_Engineer Seek_Data_Engineer Public

    Seek Answers and solutions

  2. Kafka_stack Kafka_stack Public

    Java

  3. Airflow_dags Airflow_dags Public

    Python

  4. Python_yaml_merge Python_yaml_merge Public

    This is to merge 2 yaml configurations and create a superset of the yaml file from the two configurations

    Python

  5. Springboot_prometheus Springboot_prometheus Public

    This application is to integrate spring boot with the prometheus. Simple use case is to generate a metrics on prometheus as how many requests has hit to the spring boot application

    Java

  6. Azure Azure Public