RibbitZ
Lead Data Engineer (Hands-On)
Cary, USPresencialContratoTiempo completo140.000 US$ – 145.000 US$ / año
Publicado 7 oct 2026
Este empleo está publicado en EN
Location: Cary, NC — On-site / Hybrid
Employment Type: Full-Time
Experience Required: 12–18 years
Eligibility: U.S. Citizens only
About the Engagement
This role supports the development of a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization, built on Azure Databricks.
The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and follows a modern medallion architecture (Bronze / Silver / Gold).
AI capabilities are embedded throughout the platform from day one, including ingestion, canonical mapping, data quality, reconciliation, anomaly detection, and business-user access.
This is a senior, hands-on technical leadership role. The Lead Data Engineer will own the end-to-end technical design of both the data and AI layers, establish reference implementations for the engineering team, and continue to write and review production-grade Python, Scala, and PySpark code every week.
Candidates who have not written or reviewed production code within the past year will not be considered.
Key Responsibilities
Data Platform Architecture & Engineering
Core Data Engineering
Strong hands-on experience with:
Strong experience with:
Experience establishing disciplined evaluation processes, including:
Strong experience building or working with metadata-driven frameworks involving:
Strong understanding of:
Hands-on experience with:
Experience with any of the following is highly desirable:
Please include the following information:
Employment Type: Full-Time
Experience Required: 12–18 years
Eligibility: U.S. Citizens only
About the Engagement
This role supports the development of a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization, built on Azure Databricks.
The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and follows a modern medallion architecture (Bronze / Silver / Gold).
AI capabilities are embedded throughout the platform from day one, including ingestion, canonical mapping, data quality, reconciliation, anomaly detection, and business-user access.
This is a senior, hands-on technical leadership role. The Lead Data Engineer will own the end-to-end technical design of both the data and AI layers, establish reference implementations for the engineering team, and continue to write and review production-grade Python, Scala, and PySpark code every week.
Candidates who have not written or reviewed production code within the past year will not be considered.
Key Responsibilities
Data Platform Architecture & Engineering
- Own the enterprise lakehouse architecture, including:
- Bronze / Silver / Gold data contracts
- ADLS Gen2 zone design
- Delta Lake table architecture
- Partitioning strategies
- Schema evolution
- Retention policies
- Build metadata-driven and parameterized ingestion frameworks supporting:
- Batch files
- Database extracts
- CDC feeds
- Real-time and streaming workloads
- Develop canonical PySpark and Scala Spark implementations.
- Establish engineering, coding, and automated testing standards.
- Conduct code and pull-request reviews.
- Troubleshoot production incidents and Spark workload failures.
- Optimize Spark clusters, workloads, and associated cloud costs.
- Define and enforce engineering guardrails for performance, scalability, and reliability.
- Implement CI/CD pipelines for Databricks and Azure Data Factory using:
- Azure DevOps
- Databricks Asset Bundles
- Terraform
- Design AI-assisted ingestion and canonical mapping solutions.
- Build capabilities for:
- Automated bridge-document generation
- DML generation
- Canonical table-definition generation
- AI-assisted source-to-canonical mapping
- Build AI/ML-based solutions for:
- Data drift detection
- Schema drift detection
- Volume anomaly detection
- Reconciliation failures
- Automated reconciliation
- Help design the enterprise semantic layer and knowledge graph.
- Build GPT-powered conversational data-access capabilities, including:
- Text-to-SQL
- Semantic-layer retrieval
- RAG-based access patterns
- Implement governance using Databricks Unity Catalog, including:
- Lineage
- Access control
- PII standards
- Data discovery and governance
Core Data Engineering
- 12–18 years of total experience in data engineering, data architecture, or enterprise data-platform delivery.
- Proven experience delivering enterprise-scale medallion / lakehouse architectures.
- Expert-level hands-on experience with:
- Python
- Scala
- PySpark
- Strong SQL expertise.
- Strong experience with:
- Dimensional data modeling
- Normalized data modeling
- Schema design
- Data contracts
Strong hands-on experience with:
- Delta Lake
- Unity Catalog
- Databricks Jobs & Workflows
- Cluster and pool management
- Performance tuning
- Databricks Model Serving
Strong experience with:
- ADLS Gen2
- Zone architecture
- ACLs
- Lifecycle management
- Minimum 3+ years of experience designing and deploying LLM-based systems in production.
- Strong hands-on experience with:
- RAG pipelines
- Agentic / tool-calling workflows
- Chunking strategies
- Embedding strategies
- Vector retrieval
- Hybrid retrieval
- Prompt engineering
Experience establishing disciplined evaluation processes, including:
- Golden datasets
- Regression suites
- Accuracy measurement
- Hallucination tracking
- Human-in-the-loop feedback
Strong experience building or working with metadata-driven frameworks involving:
- Schema inference
- Data profiling
- Lineage
- Data catalogs
Strong understanding of:
- Microsoft Entra ID
- Managed identities
- RBAC
- POSIX ACLs
- Azure Key Vault
- Private endpoints
- PII handling and data-security standards
Hands-on experience with:
- Azure DevOps
- Terraform
- Databricks Asset Bundles
- Automated testing for data pipelines
- Strong technical writing skills.
- Ability to communicate complex architecture decisions clearly.
- Ability to present and defend technical designs to engineers, leadership, and non-technical stakeholders.
Experience with any of the following is highly desirable:
- Knowledge graphs and ontologies
- RDF / SPARQL
- Neo4j
- Graph modeling over a lakehouse
Please include the following information:
- Updated resume
- Full Name
- Current Location
- Contact Number
- Email Address
- Work Authorization: U.S. Citizen
- LinkedIn Profile
- Availability to Start
- Interview Availability for the next 3 days
Resumen del puesto
Tipo de empleo
Tiempo completo
Correo electrónico
business@ribbitzllc.com
Habilidades requeridas
PythonScalaPySparkSpark workload troubleshooting and performance tuningDelta Lake and lakehouse architecture (Bronze/Silver/Gold, partitioning, schema evolution, retention)Databricks (Jobs & Workflows, Model Serving, Asset Bundles, cluster and pool management)ADLS Gen2 and Azure data stack (zone architecture, ACLs, lifecycle management)Streaming and ingestion frameworks (Kafka, Azure Event Hubs, CDC, Spark Structured Streaming)SQL and data modeling (dimensional and normalized modeling, schema design, data contracts)CI/CD and Infrastructure as Code (Azure DevOps, Terraform, automated testing for data pipelines)Generative AI / LLM engineering (RAG, embeddings, chunking, LangChain/LlamaIndex/agentic workflows, prompt engineering)AI/ML-based data quality and reconciliation (data drift, schema drift, anomaly detection, synthetic test data)Metadata and data governance (schema inference, profiling, lineage, data catalogs, Unity Catalog)Azure security and identity (Microsoft Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling)Technical leadership and communication (mentoring, code reviews, architecture reviews, presenting to technical and non-technical stakeholders)
Empleos similares
Senior Product Owner – ERP Finance Transformation
I8IS INC.
Cary, USPresencialContratoTiempo completo
hace 17 días
API Architect – Azure
I8IS INC.
Cary, USPresencialContratoTiempo completo
hace 19 días
Senior Product Manager
OTSI
Cary, USPresencialContratoTiempo completo
hace 23 días
Internal Medicine Job Near Cary, NC
Atlantic MEDsearch
Cary, USPresencialContratoTiempo completo
hace 14 días
Family Practice Job Near Cary, NC
Atlantic MEDsearch
Cary, USPresencialContratoTiempo completo
hace 15 días
Site Reliability Engineer (SRE) – Security Infrastructure- Hybrid
Simple Solutions
Cary, USPresencialContratoTiempo completo
hace 23 días