About Me
Metric-driven Data Scientist and Engineering student with hands-on experience building complex data pipelines and predictive ML models. Successfully engineered robust Python OCR pipelines achieving 92% accuracy on unstructured technical data, and deployed a Two-Stage Machine Learning classification system for financial analytics. Highly proficient in Pandas, SQL, and translating messy, unstructured datasets into actionable business insights.
Skills & Technologies
Work Experience
Artificial Intelligence and Digital Data Intern
Contributed to document intelligence projects by developing solutions for extracting structured data from PDFs. Created OCR data pipelines and automated workflows to streamline the processing of technical documents. Achieved 4-36x faster processing speeds with a custom bilingual OCR data pipeline and ensured over 92% average accuracy on structured engineering tables.
Projects
End-to-End Cancer Risk Prediction System using Optuna-Tuned XGBoost
Built a multi-class cancer risk stratification pipeline on a 2,000-patient dataset with 18 clinical features, achieving 88% accuracy and 0.72 macro-F1 using class-weighted XGBoost optimized via Optuna Bayesian hyperparameter tuning. Deployed an interactive Streamlit web app with batch and single-patient prediction modes.
Two-Stage Loan Approval & Amount Prediction System
Engineered a two-stage ML pipeline combining a Random Forest Classifier for loan approval prediction and a Random Forest Regressor for loan amount estimation, trained on 4,269 applicant records. Optimized models via GridSearchCV with 5-fold cross-validation, achieving a classifier OOB score of 98.4%.
AI-Powered Customer Support Agent with Memory
Engineered a full-stack AI insurance claims copilot using FastAPI, LangChain, and a Streamlit adjuster workbench, automating coverage recommendation generation for FNOL submissions. Designed a three-layer context retrieval pipeline.
Education
BTECH in Computer Science
Data Science