
Naveen M
Kompetenzen

Meine Dienstleistungen


Arbeitserfahrung
Data Engineer
NeuroDiscovery AI • Vollzeit
Nov 2025 - Present • 11 mos
NeuroDiscovery AI | Gurugram, Haryana | Nov 2025 - Present • Architected and built the company’s core healthcare data platform from scratch, supporting analytics, internal tooling and AI driven product capabilities. • Designed scalable ETL pipelines using PySpark on Azure Databricks to ingest and process healthcare datasets covering 5M+ patient records and 10+ TB of clinical data, including physician progress notes, EHR records and medical documentation. • Built data ingestion pipelines sourcing healthcare records from major EHR systems: eClinicalWorks (eCW), athenaOne, athenaPractice and Greenway Health. • Led onboarding and normalization of 5M+ patient records from multiple healthcare providers, mapping heterogeneous EHR schemas into a common data model and reconciling medical coding systems (ICD-10, CPT, SNOMED, LOINC). • Designed document processing pipelines converting unstructured clinical progress notes into structured medical sections — Review of Systems (ROS), Past Medical History, Observations and Clinical Assessments. • Built a scalable clinical data warehouse for progress-notes datasets on Databricks, enabling structured querying and analytics across large medical document repositories. • De-identify protected health information (PHI) in line with HIPAA requirements before clinical data enters the analytics environment. • Orchestrate scheduled workflows in Apache Airflow and manage transformation models in dbt, with validation and data quality checks that catch schema drift before it reaches downstream consumers.