About Me
Data Engineer with 2+ years of experience building production ETL and data pipelines on AWS and GCP using SQL, PySpark, and Python. Hands-on experience with AWS Glue, Apache Airflow, Amazon S3, AWS DMS, Kinesis, Snowflake, BigQuery, and Cloud Composer. Experienced in CDC-based ingestion, Data Lakehouse architecture, data quality validation, incremental processing, and Spark performance optimization.
Skills & Technologies
Work Experience
Data Engineer
Working as a Data Engineer at Sify Technologies designing and implementing production ETL pipelines on AWS and GCP. Built data lakehouse pipelines using AWS Glue, PySpark, Kinesis, DMS, Airflow, Snowflake, and GCP tools. Focused on scalable pipeline orchestration, data quality validation, and performance optimization for large data volumes.
Projects
Shoppers Stop Data Lakehouse
Designed and implemented end-to-end data pipelines on AWS using Kinesis, AWS DMS CDC, AWS Glue, S3, PySpark, and Airflow following a Medallion architecture for scalable and maintainable data delivery. Built PySpark and SQL transformations for data cleansing, validation, and enrichment. Integrated AWS S3 with Snowflake for downstream analytics.
Network Telemetry Data Analytics Engine
Developed batch data pipeline on GCP to ingest, clean, and load telemetry logs from GCS and Microsoft SQL Server into BigQuery. Processed 50–100 GB daily within 2–3 hour windows. Utilized SQL and Python to clean, validate and deduplicate data handling dynamic schema changes. Improved query performance over 35% by partitioning and clustering BigQuery tables.
Education
B.Tech
Computer Engineering