● Open to Work

chaitaya

Data Scientist
Bengaluru, India

About Me

Data Scientist with expertise in Python, SQL, Machine Learning, NLP, and GenAI. Skilled in data preprocessing, exploratory data analysis, and predictive modeling using Pandas, NumPy, and Scikit-learn, with hands-on experience building and deploying end-to-end ML and RAG-based systems that turn data into actionable, decision-ready insights.

Skills & Technologies

PythonSQLMachine LearningNLPClassificationRegressionPredictive ModelingDemand ForecastingFeature EngineeringModel TrainingPandasNumPyScikit-learnXGBoostLightGBMMatplotlibSeabornSciPyOpenAI APIHugging FaceRAGFAISSChromaDBMLflowDockerGitHub ActionsFastAPITF-IDFLogistic RegressionSVMHOG feature extractionImage preprocessingCross-validationPower BITableauStreamlitAdvanced ExcelJupyter NotebookGitPostgreSQL

Work Experience

Data Science Intern

AI Variant
Sep 2025Mar 2026

Performed data cleaning, preprocessing, and exploratory data analysis on over 5,000 records using Python, Pandas, and NumPy to prepare datasets for ML development. Built classification and regression models and developed interactive Power BI dashboards, translating KPIs and model outputs into stakeholder-ready insights.

Data Science Intern

IntrnForte
Sep 2024Jan 2025

Optimized preprocessing pipelines for structured and unstructured datasets reducing processing time by 20%. Built classification models using Scikit-learn, Pandas, and NumPy achieving up to 88% prediction accuracy.

Projects

Supplier Stock Prediction & Inventory Optimization System

Built an end-to-end AI-powered inventory system forecasting product demand and flagging stockout/overstock risk to support data-driven supplier ordering decisions. Engineered demand-trend, seasonality, and lead-time features from historical sales, inventory, and supplier data in PostgreSQL; trained XGBoost/LightGBM models for forecasting plus a risk-classification layer for stockout/overstock alerts. Deployed predictions via a FastAPI backend with a Streamlit dashboard, tracked experiments in MLflow, containerized with Docker, and automated deployment through GitHub Actions CI/CD. Integrated an LLM-based RAG layer (OpenAI/Hugging Face, FAISS/ChromaDB) enabling natural-language queries over forecasts, inventory, and supplier data.

PythonSQLXGBoostLightGBM+9

Sentiment Analysis of Restaurant Reviews

Built an NLP sentiment-classification model using text preprocessing and TF-IDF feature extraction with Logistic Regression. Achieved 87% accuracy on 5,000+ reviews, enabling automated sentiment insights for data-driven decision-making.

PythonNLTKScikit-learnLogistic Regression+1

Neonatal Jaundice Detection

Built a machine learning model to assist early detection of neonatal jaundice from clinical datasets using image preprocessing, HOG feature extraction, and SVM classification. Improved prediction reliability and reduced false positives, supporting more accurate medical screening. Preprocessed and augmented a labeled clinical image dataset applying grayscale conversion, resizing, and normalization. Tuned SVM hyperparameters and validated using cross-validation, benchmarking performance against baseline classifiers.

PythonScikit-learnSVMJupyter Notebook+3

Education

Bachelor of Engineering

Computer Science and Engineering

MVJ College of Engineering, Bengaluru2021 - 2025
chaitaya - Data Scientist | HiringAnt