4+ years across mobility, enterprise, and healthcare · Data Scientist @ Uber · M.S. Data Science @ George Washington University
Data Scientist with 4+ years building forecasting, experimentation, and analytics systems across mobility, enterprise, and healthcare domains. Currently at Uber, working on demand forecasting, pricing experiments, and automated reporting pipelines for marketplace and rider analytics.
Previously designed enterprise forecasting models and churn analytics at Cognizant, and began in healthcare at Cipla, where I handled commercial reporting, inventory forecasting, and data quality across pharmaceutical business units.
I work end to end — from ingestion and feature engineering through model evaluation and deployment — primarily in Python, SQL, PySpark, Databricks, and Snowflake. Completed an M.S. in Data Science at George Washington University in December 2025.
- Processed 3M+ marketplace records through PySpark and Databricks workflows supporting demand forecasting across Uber mobility teams
- Cut forecasting RMSE from 197 to 137 by benchmarking XGBoost, GRU, and LSTM under rolling-window validation across 177K+ records
- Raised rare-entity detection by 18 points to 95.07% accuracy through tokenization optimization and weighted loss tuning on transformer models
- Reduced turnaround on 100+ recurring analytics requests by automating pipelines with FastAPI, Snowflake, and Kafka
- Validated 1M+ healthcare records, resolving reporting discrepancies and improving accuracy across BI and planning systems
| Project | Description | Stack |
|---|---|---|
| NLP Entity Classification | Fine-tuned BERT and ELECTRA to automate PII entity classification across 20K+ student essay records. Tokenization optimization and weighted loss tuning raised rare-entity detection by 18 points, reaching 95.07% accuracy. | PyTorch · Hugging Face · BERT · ELECTRA |
| Cost Forecasting Pipeline | Forecasting meal production costs across 177K+ records from 100+ schools. Benchmarked Linear Regression, XGBoost, GRU, and LSTM under rolling-window validation, cutting RMSE from 197 to 137. | PyTorch · scikit-learn · XGBoost · LSTM/GRU |
| LA Crime & Geospatial Analysis | Spatial and temporal analysis of Los Angeles crime data from 2020 onward, identifying trends to inform law enforcement and policy stakeholders. | Python · pandas · geospatial |
- Built PySpark, SQL, and Databricks analytics workflows processing 3M+ marketplace and rider interaction records for demand forecasting and operational reporting
- Engineered XGBoost and SQL modeling workflows across 4+ pricing experiments, identifying rider engagement and conversion patterns
- Developed automated data pipelines using FastAPI, Snowflake, and Kafka, reducing turnaround for 100+ recurring analytics requests from product stakeholders
- Designed Python, scikit-learn, and SQL forecasting models across 7 enterprise workflows supporting customer analytics and operational reporting
- Streamlined PySpark and BigQuery ETL pipelines processing 2M+ transactional records, improving KPI reporting efficiency
- Performed customer segmentation and churn analysis with XGBoost and feature engineering across 10+ retail datasets
- Automated recurring analytics and model scoring on AWS across 300K+ CRM records
- Analyzed healthcare and commercial datasets supporting forecasting and inventory planning across 6 pharmaceutical business units
- Built Tableau dashboards tracking sales performance and regional distribution across 8 therapeutic categories
- Investigated reporting discrepancies within 1M+ healthcare records using SQL and data-quality validation
- Automated recurring reporting workflows serving 25+ operational requests across analytics, supply chain, and commercial teams
M.S. Data Science · George Washington University · Jan 2024 – Dec 2025
Arlington, Virginia · varshithreddy3003@gmail.com · linkedin.com/in/varshithreddy77
Open to Data Scientist, ML Engineer, Data Engineer, and Product Analytics roles.
