s
samyakjain245

Samyak Jain

@samyakjain245

Data Engineer

Indien
Englisch, Hindi
Einige Informationen werden in englischer Sprache angezeigt.
Über mich
Hi! I’m a Data Engineer with 3.5+ years of experience building reliable data pipelines, ETL processes, and scalable data solutions using Python, SQL, PySpark, Databricks, and Microsoft Azure. I can help with: • Data cleaning & transformation • ETL/ELT pipelines • PySpark & Databricks • SQL & database solutions • Azure ADLS & cloud pipelines • Data quality & Power BI dashboards I focus on clean, maintainable code, accurate results, and reliable delivery.... Mehr lesen

Kompetenzen

s
samyakjain245
Samyak Jain
offline • 
Durchschnittliche Antwortzeit: 1 Stunde

Meine Dienstleistungen

Beratung im Bereich Datentechnik
I will develop etl,elt pipeline using azure databricks and pyspark

Arbeitserfahrung

ZS

ETL Developer

ZS • Vollzeit

Feb 2024 - Present2 yrs 6 mos

• Architected, Model and owned large-scale ETL/ELT pipelines using SQL and Spark (PySpark) application to process Large-volumes healthcare and Pharma data, for analytics and reporting. • Built and supported executive and market intelligence dashboards focused on pharmacy marketing, brand performance, market share, and regional sales trends, enabling data flow, data architectures and data-driven decisions. • Orchestrated and monitored Spark workflows on EMR and Hive compatible environments using Azkaban, AWS EMR, and Linux/Putty, driving scalable and reliable data processing. • Optimized(optimization) and Automation ETL and Spark workloads by allocating shuffle partitions, executor memory/ cores, and driver configurations, improving processing speed and overall pipeline stability. • Collaborated with cross-functional teams to streamline Fintech, SD, HUB, and pharmacy marketing analytics workflows, accelerating data ingestion, quality assurance, Security and enabling self- service analytics. • Implement Master Data Management (MDM) and data governance frameworks for products, brands, territories, and customer hierarchies, ensuring data quality, consistency, and trusted executive reporting. • Partnered with analysts, Bl engineers, and product managers to support statistical analysis and downstream machine learning use cases.

Celebal_Technologies

Data Engineer

Celebal Technologies • Vollzeit

Dec 2022 - Dec 20231 yr

Engineered scalable data pipelines and ETL pipelines using SQL, Python, Spark (PySpark/Spark SQL) on Azure Databricks, data integration from databases, cloud storage, S3 and SharePoint(Data Ingestion formats (Parquet, JSON, Iceberg, CSV)). • Designed and implemented data enrichments and transformations using Medallion Architecture (Bronze/Silver/Gold) with Delta Lake to support analytical and reporting use cases. • Developed a data quality audit engine and integrated datasets with Power BI to enable reliable visual analytics and business insights. • Proficient in Microsoft Azure for building and operating scalable data engineering solutions, leveraging cloud storage and distributed processing. • Created Delta Live Table pipelines for incremental and real-time data processing, ensuring continuous updates across bronze, silver, and gold layers. • Developed advanced SQL procedures, Database Management, Database Design and reusable transformation scripts for optimized data movement and manufacturing concepts.