Location: New Jersey - US On site Hybrid
Experience: 12+ Years
C2C W2 Full time
Visa – USC/ GC/H4EAD/GC EAD (NO H1B)
Rate $65/hr C2C
Linkedin ID
Passport number Mandatory for GC GC EAD H4EAD
Job Description –
Job Summary
We are seeking a highly experienced Lead Data Engineer (10+ years) with deep expertise in Snowflake, DBT, Apache Airflow, and StreamSets, and strong hands-on experience in designing
enterprise-grade ETL/ELT, data migration, and multi-source ingestion frameworks within the Life Sciences domain.
This role will lead large-scale data platform modernization initiatives including legacy-to-cloud migrations, cross-system integrations, and enterprise data harmonization in regulated environments.
Key Responsibilities
1. Snowflake Architecture & Enterprise Data Platform Design
Lead architecture and implementation of scalable Snowflake data platforms:
o Multi-layered architecture (Landing → Raw → Staging → Curated → Data Marts)
o Separation of compute and storage optimization
o Multi-cluster warehouses and workload isolation
Design secure cross-account data sharing strategies.
Implement:
o Snowpipe for automated ingestion
o Streams & Tasks for CDC-based incremental processing
o Time Travel & Zero-copy cloning for environment management
Implement data masking, row-level security, and RBAC frameworks.
Optimize storage, partitioning (micro-partition pruning), and query performance.
2. Data Migration & Modernization
Lead end-to-end data migration initiatives including:
o Legacy data warehouse (Teradata, Oracle, SQL Server, Netezza) to Snowflake
o On-prem to cloud modernization programs
Conduct:
o Source system analysis and profiling
o Data quality assessment and remediation planning
o Schema conversion and transformation mapping
Design migration frameworks:
o Bulk historical data loads
o Incremental migration strategies
o Parallel-run validation strategies
Perform reconciliation and data validation between legacy and target systems.
Develop automated validation scripts using SQL and DBT tests.
Support cutover planning and production readiness.
3. Data Ingestion & Multi-Source Integration
Design and implement ingestion frameworks for structured, semi-structured, and unstructured data
from multiple enterprise systems:
Structured Sources
Oracle, SQL Server, SAP, PostgreSQL
Clinical systems (EDC, CDMS, CTMS)
Regulatory systems (RIM)
Commercial systems (CRM, ERP)
Semi-Structured Sources
JSON, XML, Avro files
API responses
External vendor feeds
Unstructured Sources (where applicable)
Document metadata ingestion
Log and audit trail ingestion
Ingestion Responsibilities
Build ingestion pipelines using:
o StreamSets for batch and streaming ingestion
o Snowpipe with cloud storage integration (S3/Azure Blob/GCS)
o API-driven ingestion frameworks
Implement CDC mechanisms using:
o Database log-based CDC
o Timestamp-based incremental extraction
o Snowflake Streams
Develop metadata-driven ingestion frameworks.
Design resilient pipelines with error handling, retry logic, and monitoring.
Ensure schema evolution handling and version control.
4. DBT – Enterprise Transformation Framework
Architect and govern DBT transformation layers:
o Staging models
o Intermediate models
o Data marts
Implement:
o Incremental models
o Snapshot strategies for historical tracking
o Surrogate key management
Develop custom macros and reusable transformation components.
Implement comprehensive DBT testing framework:
o Source freshness tests
o Schema validation tests
o Business rule validation
Generate lineage documentation for audit and regulatory needs.
Optimize DBT models specifically for Snowflake compute efficiency.
5. ETL / ELT Orchestration & Automation
Design ELT-first architecture leveraging Snowflake processing power.
Orchestrate complex workflows using Apache Airflow:
o DAG dependency management
o SLA monitoring
o Automated recovery workflows
Implement CI/CD for:
o DBT deployments
o Airflow pipelines
o Snowflake objects
Build data observability frameworks (pipeline monitoring, anomaly detection).
6. Enterprise Data Modelling
Design scalable data models:
o Dimensional (Star/Snowflake schemas)
o Data Vault 2.0 (for auditability and traceability)
o Canonical data models
Align models with Life Sciences business domains:
o Clinical trial lifecycle
o Regulatory submissions
o Pharmacovigilance
o Commercial analytics
Support cross-domain data harmonization.
7. Life Sciences Domain Expertise
Experience delivering data platforms supporting:
Clinical trial data (EDC, CDMS, CTMS)
Regulatory and submission systems
Pharmacovigilance & safety systems
Commercial & sales analytics
Real-World Evidence (RWE)
Ensure compliance with:
GxP validation standards
21 CFR Part 11
HIPAA / GDPR
ALCOA+ principles
Support audit readiness and regulatory traceability.
Required Qualifications
10+ years of experience in Data Engineering and Enterprise Data Platforms.
4–6+ years hands-on Snowflake implementation experience.
Strong experience in:
o Large-scale data migration programs
o Multi-source data ingestion frameworks
o DBT advanced transformation design
o Apache Airflow orchestration
o StreamSets ingestion pipelines
Advanced SQL expertise.
Experience in Life Sciences domain projects.
Cloud platform experience (AWS/Azure/GCP).
Regards
| ||