Data Analyst Β· AI/ML Engineer Β· Data Scientist Β· Business analyst
I build data systems that are honest about their limits and useful in production.
I work at the intersection of data engineering, machine learning, and business intelligence β turning messy, large-scale datasets into systems that stakeholders can actually trust.
- ποΈ Data pipelines that handle the dirty work β malformed HTML, schema drift, 60GB+ of raw filings
- π Dashboards and KPI systems built for non-technical consumers, not just engineers
- π€ ML pipelines where integrity matters more than headline accuracy
- π Automated retraining workflows that stay accurate without manual intervention
Architected a Polars-based out-of-core pipeline for 60GB+ SEC EDGAR filings on commodity hardware β Corporate insolvency prediction 12 months ahead
Medallion architecture (Bronze β Silver β Gold) | Polars Lazy Evaluation | FinBERT MD&A sentiment | XGBoost production model
- 4.5Γ lift over the 8.4% baseline crash rate
- XGBoost: 87% recall, 38% precision, F1 0.52 | 1,786 false positives vs LSTM's 24,652
- Feature Order Lock prevents silent prediction drift at inference time
- π΄ Live Demo β HuggingFace Spaces
Polars Parquet FinBERT XGBoost SHAP DuckDB Streamlit BeautifulSoup
Stochastic financial runway modeling for gig economy income volatility
Full-stack forecasting platform | Hybrid stacking ensemble | Nightly automated retraining
- XGBoost (0.6) + Random Forest (0.4) ensemble generating bounded planning ranges from rolling forecast-error variance
- SHA-256 fingerprinting for immutable data lineage across retraining cycles
- Sub-200ms query response via Supabase SECURITY DEFINER views
- π΄ Live Demo β Vercel
FastAPI XGBoost Next.js 14 TypeScript Supabase PostgreSQL GitHub Actions
π·οΈ Loan Approval Prediction
I reduced model accuracy from 98% to 88% β and that was the win.
Found a data leakage flaw (pre-split oversampling β synthetic duplicates bleeding into test set). Rebuilt from scratch.
- 11 classifiers Γ 100 randomised partitions β macro F1 0.81, rejected-class F1 0.71 on leakage-proof holdout
- CatBoost selected for consistency, not peak score
- Feature importance: Credit History ~24%, Loan Amount ~19%, Applicant Income ~18%
Scikit-Learn CatBoost XGBoost imbalanced-learn Pandas Seaborn
π Veri-Vigil AI β ET GenAI Hackathon 2026
Browser-based content trust analyzer | ET GenAI Hackathon Semi-Finalist
Chrome Extension (Manifest V3) that analyses YouTube video metadata and generates a trust score + explanation using LLaMA 3.1 via Groq API.
FastAPI LLaMA 3.1 Groq Chrome Extension Manifest V3
Languages & Querying
ML & Data Engineering
Visualization & BI
Infrastructure & DevOps
- π― OJEE 2023 β Top 5% Merit rank (800 / 16,000+ candidates)
- π₯ ET GenAI Hackathon 2026 β Semi-Finalist | Built Veri-Vigil AI in 48 hours
- π Trithon 2023 β Cash Prize winner for problem-solving and technical collaboration
- π RAECC-2025 National Conference β Presented research on E-Waste upcycling into Biodegradable 3D Printing Filaments
- π B.Tech CSE (AI) β GIFT Autonomous, Bhubaneswar | Graduating in June 2026
I'm actively looking for Data Analyst, AI/ML Engineer, and Business Intelligence roles β available full-time from June 2026, open to relocation.
"The model that admits its flaws is the one you can trust in production."