Skip to content
View varshithreddy77's full-sized avatar
🎯
Open to Work
🎯
Open to Work

Highlights

  • Pro

Block or report varshithreddy77

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
varshithreddy77/README.md

Varshith Reddy Bhimireddy

Data Scientist · Forecasting & Experimentation · Large-Scale Analytics

LinkedIn GitHub Email

4+ years across mobility, enterprise, and healthcare · Data Scientist @ Uber · M.S. Data Science @ George Washington University


Expertise Areas

Time Series Forecasting Experimentation Causal Inference Product Analytics NLP Distributed Computing MLOps Data Pipelines BI


About

Data Scientist with 4+ years building forecasting, experimentation, and analytics systems across mobility, enterprise, and healthcare domains. Currently at Uber, working on demand forecasting, pricing experiments, and automated reporting pipelines for marketplace and rider analytics.

Previously designed enterprise forecasting models and churn analytics at Cognizant, and began in healthcare at Cipla, where I handled commercial reporting, inventory forecasting, and data quality across pharmaceutical business units.

I work end to end — from ingestion and feature engineering through model evaluation and deployment — primarily in Python, SQL, PySpark, Databricks, and Snowflake. Completed an M.S. in Data Science at George Washington University in December 2025.


Key Achievements

MS Data Science Google Data Analytics Advanced Data Analytics

  • Processed 3M+ marketplace records through PySpark and Databricks workflows supporting demand forecasting across Uber mobility teams
  • Cut forecasting RMSE from 197 to 137 by benchmarking XGBoost, GRU, and LSTM under rolling-window validation across 177K+ records
  • Raised rare-entity detection by 18 points to 95.07% accuracy through tokenization optimization and weighted loss tuning on transformer models
  • Reduced turnaround on 100+ recurring analytics requests by automating pipelines with FastAPI, Snowflake, and Kafka
  • Validated 1M+ healthcare records, resolving reporting discrepancies and improving accuracy across BI and planning systems

Featured Projects

Project Description Stack
NLP Entity Classification Fine-tuned BERT and ELECTRA to automate PII entity classification across 20K+ student essay records. Tokenization optimization and weighted loss tuning raised rare-entity detection by 18 points, reaching 95.07% accuracy. PyTorch · Hugging Face · BERT · ELECTRA
Cost Forecasting Pipeline Forecasting meal production costs across 177K+ records from 100+ schools. Benchmarked Linear Regression, XGBoost, GRU, and LSTM under rolling-window validation, cutting RMSE from 197 to 137. PyTorch · scikit-learn · XGBoost · LSTM/GRU
LA Crime & Geospatial Analysis Spatial and temporal analysis of Los Angeles crime data from 2020 onward, identifying trends to inform law enforcement and policy stakeholders. Python · pandas · geospatial

Professional Experience

Data Scientist · Uber · Feb 2026 – Present

  • Built PySpark, SQL, and Databricks analytics workflows processing 3M+ marketplace and rider interaction records for demand forecasting and operational reporting
  • Engineered XGBoost and SQL modeling workflows across 4+ pricing experiments, identifying rider engagement and conversion patterns
  • Developed automated data pipelines using FastAPI, Snowflake, and Kafka, reducing turnaround for 100+ recurring analytics requests from product stakeholders

Data Scientist · Cognizant · Feb 2023 – Dec 2023

  • Designed Python, scikit-learn, and SQL forecasting models across 7 enterprise workflows supporting customer analytics and operational reporting
  • Streamlined PySpark and BigQuery ETL pipelines processing 2M+ transactional records, improving KPI reporting efficiency
  • Performed customer segmentation and churn analysis with XGBoost and feature engineering across 10+ retail datasets
  • Automated recurring analytics and model scoring on AWS across 300K+ CRM records

Data Analyst · Cipla · Jul 2020 – Feb 2023

  • Analyzed healthcare and commercial datasets supporting forecasting and inventory planning across 6 pharmaceutical business units
  • Built Tableau dashboards tracking sales performance and regional distribution across 8 therapeutic categories
  • Investigated reporting discrepancies within 1M+ healthcare records using SQL and data-quality validation
  • Automated recurring reporting workflows serving 25+ operational requests across analytics, supply chain, and commercial teams

Technology Stack

Languages & Query

Python SQL PySpark R

Machine Learning & Statistics

scikit-learn XGBoost PyTorch TensorFlow Hugging Face SHAP

Data Engineering & Big Data

Spark Databricks Airflow Kafka MLflow FastAPI Docker

Cloud & Warehouses

AWS Snowflake BigQuery GCP

Databases

PostgreSQL MySQL MongoDB

Visualization & Reporting

Tableau Power BI Looker


Education

M.S. Data Science · George Washington University · Jan 2024 – Dec 2025


Arlington, Virginia · varshithreddy3003@gmail.com · linkedin.com/in/varshithreddy77

Open to Data Scientist, ML Engineer, Data Engineer, and Product Analytics roles.

Pinned Loading

  1. cost-forecasting-ml-deep-learning cost-forecasting-ml-deep-learning Public template

    Forked from Areena2908/fall-2025-group9

    HTML

  2. NLP-entity-classification-transformers NLP-entity-classification-transformers Public

    Jupyter Notebook

  3. Los-Angeles-Crime-and-Geospatial-Trend-Analysis Los-Angeles-Crime-and-Geospatial-Trend-Analysis Public

    Forked from varshith2233/DATS6103_Team1

    This project analyzes Los Angeles crime data from 2020 to present, aiming to identify trends and inform stakeholders, including law enforcement and policymakers, to enhance public safety measures. …

    Python