About Me
Data Scientist with expertise in Python, SQL, Machine Learning, NLP, and GenAI. Skilled in data preprocessing, exploratory data analysis, and predictive modeling using Pandas, NumPy, and Scikit-learn, with hands-on experience building and deploying end-to-end ML and RAG-based systems that turn data into actionable, decision-ready insights.
Skills & Technologies
Work Experience
Data Science Intern
Performed data cleaning, preprocessing, and exploratory data analysis on over 5,000 records using Python, Pandas, and NumPy to prepare datasets for ML development. Built classification and regression models and developed interactive Power BI dashboards, translating KPIs and model outputs into stakeholder-ready insights.
Data Science Intern
Optimized preprocessing pipelines for structured and unstructured datasets reducing processing time by 20%. Built classification models using Scikit-learn, Pandas, and NumPy achieving up to 88% prediction accuracy.
Projects
Supplier Stock Prediction & Inventory Optimization System
Built an end-to-end AI-powered inventory system forecasting product demand and flagging stockout/overstock risk to support data-driven supplier ordering decisions. Engineered demand-trend, seasonality, and lead-time features from historical sales, inventory, and supplier data in PostgreSQL; trained XGBoost/LightGBM models for forecasting plus a risk-classification layer for stockout/overstock alerts. Deployed predictions via a FastAPI backend with a Streamlit dashboard, tracked experiments in MLflow, containerized with Docker, and automated deployment through GitHub Actions CI/CD. Integrated an LLM-based RAG layer (OpenAI/Hugging Face, FAISS/ChromaDB) enabling natural-language queries over forecasts, inventory, and supplier data.
Sentiment Analysis of Restaurant Reviews
Built an NLP sentiment-classification model using text preprocessing and TF-IDF feature extraction with Logistic Regression. Achieved 87% accuracy on 5,000+ reviews, enabling automated sentiment insights for data-driven decision-making.
Neonatal Jaundice Detection
Built a machine learning model to assist early detection of neonatal jaundice from clinical datasets using image preprocessing, HOG feature extraction, and SVM classification. Improved prediction reliability and reduced false positives, supporting more accurate medical screening. Preprocessed and augmented a labeled clinical image dataset applying grayscale conversion, resizing, and normalization. Tuned SVM hyperparameters and validated using cross-validation, benchmarking performance against baseline classifiers.
Education
Bachelor of Engineering
Computer Science and Engineering