Volver a empleos
ScovaiScovaiJobs
RibbitZ

RibbitZ

Lead Data Engineer (Hands-On)

Cary, USPresencialContratoTiempo completo140.000 US$ – 145.000 US$ / año

Publicado 7 oct 2026

Este empleo está publicado en EN

Location: Cary, NC — On-site / Hybrid
Employment Type: Full-Time
Experience Required: 12–18 years
Eligibility: U.S. Citizens only
About the Engagement
This role supports the development of a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization, built on Azure Databricks.
The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and follows a modern medallion architecture (Bronze / Silver / Gold).
AI capabilities are embedded throughout the platform from day one, including ingestion, canonical mapping, data quality, reconciliation, anomaly detection, and business-user access.
This is a senior, hands-on technical leadership role. The Lead Data Engineer will own the end-to-end technical design of both the data and AI layers, establish reference implementations for the engineering team, and continue to write and review production-grade Python, Scala, and PySpark code every week.
Candidates who have not written or reviewed production code within the past year will not be considered.
Key Responsibilities
Data Platform Architecture & Engineering
  • Own the enterprise lakehouse architecture, including:
  • Bronze / Silver / Gold data contracts
  • ADLS Gen2 zone design
  • Delta Lake table architecture
  • Partitioning strategies
  • Schema evolution
  • Retention policies
Design and implement scalable, production-grade data engineering patterns for enterprise workloads.Ingestion & Streaming Frameworks
  • Build metadata-driven and parameterized ingestion frameworks supporting:
  • Batch files
  • Database extracts
  • CDC feeds
  • Real-time and streaming workloads
Work extensively with:Azure Event HubsKafkaSpark Structured StreamingHands-On Engineering Leadership
  • Develop canonical PySpark and Scala Spark implementations.
  • Establish engineering, coding, and automated testing standards.
  • Conduct code and pull-request reviews.
  • Troubleshoot production incidents and Spark workload failures.
  • Optimize Spark clusters, workloads, and associated cloud costs.
  • Define and enforce engineering guardrails for performance, scalability, and reliability.
CI/CD, Infrastructure & Observability
  • Implement CI/CD pipelines for Databricks and Azure Data Factory using:
  • Azure DevOps
  • Databricks Asset Bundles
  • Terraform
Establish observability using:Azure MonitorLog AnalyticsDrive automated testing and repeatable deployment practices across environments.AI-Augmented Data Engineering
  • Design AI-assisted ingestion and canonical mapping solutions.
  • Build capabilities for:
  • Automated bridge-document generation
  • DML generation
  • Canonical table-definition generation
  • AI-assisted source-to-canonical mapping
Implement appropriate human-review and approval gates for AI-generated mappings and transformations.AI-Driven Data Quality & Reconciliation
  • Build AI/ML-based solutions for:
  • Data drift detection
  • Schema drift detection
  • Volume anomaly detection
  • Reconciliation failures
  • Automated reconciliation
Develop privacy-preserving synthetic test-data approaches.Semantic Layer & Generative AI
  • Help design the enterprise semantic layer and knowledge graph.
  • Build GPT-powered conversational data-access capabilities, including:
  • Text-to-SQL
  • Semantic-layer retrieval
  • RAG-based access patterns
Enforce row-level and column-level security within AI-powered data-access solutions.Governance & Technical Leadership
  • Implement governance using Databricks Unity Catalog, including:
  • Lineage
  • Access control
  • PII standards
  • Data discovery and governance
Participate in and lead:Architecture Review BoardsAI governance forumsDesign reviewsMentor engineers and establish strong technical documentation standards.Present and defend architecture decisions to both technical and non-technical stakeholders.Must-Have Skills & Experience
Core Data Engineering
  • 12–18 years of total experience in data engineering, data architecture, or enterprise data-platform delivery.
  • Proven experience delivering enterprise-scale medallion / lakehouse architectures.
  • Expert-level hands-on experience with:
  • Python
  • Scala
  • PySpark
Strong ability to build modular, well-tested, production-grade data solutions.Deep experience troubleshooting and tuning Spark workloads.Experience optimizing large-scale batch and streaming pipelines using Delta Lake.SQL & Data Modeling
  • Strong SQL expertise.
  • Strong experience with:
  • Dimensional data modeling
  • Normalized data modeling
  • Schema design
  • Data contracts
Databricks
Strong hands-on experience with:
  • Delta Lake
  • Unity Catalog
  • Databricks Jobs & Workflows
  • Cluster and pool management
  • Performance tuning
  • Databricks Model Serving
Microsoft Azure Data Stack
Strong experience with:
  • ADLS Gen2
  • Zone architecture
  • ACLs
  • Lifecycle management
Azure Data FactoryParameterized pipelinesMetadata-driven ingestion frameworksAzure Event HubsGenerative AI / LLM Engineering
  • Minimum 3+ years of experience designing and deploying LLM-based systems in production.
  • Strong hands-on experience with:
  • RAG pipelines
  • Agentic / tool-calling workflows
  • Chunking strategies
  • Embedding strategies
  • Vector retrieval
  • Hybrid retrieval
  • Prompt engineering
Experience with at least one of:LangChainLlamaIndexLangGraphExperience with at least one production AI stack:Azure OpenAIOpenAIDatabricks Model ServingLLM Evaluation & Quality
Experience establishing disciplined evaluation processes, including:
  • Golden datasets
  • Regression suites
  • Accuracy measurement
  • Hallucination tracking
  • Human-in-the-loop feedback
Metadata & Governance
Strong experience building or working with metadata-driven frameworks involving:
  • Schema inference
  • Data profiling
  • Lineage
  • Data catalogs
Azure Security
Strong understanding of:
  • Microsoft Entra ID
  • Managed identities
  • RBAC
  • POSIX ACLs
  • Azure Key Vault
  • Private endpoints
  • PII handling and data-security standards
DevOps & Infrastructure as Code
Hands-on experience with:
  • Azure DevOps
  • Terraform
  • Databricks Asset Bundles
  • Automated testing for data pipelines
Communication
  • Strong technical writing skills.
  • Ability to communicate complex architecture decisions clearly.
  • Ability to present and defend technical designs to engineers, leadership, and non-technical stakeholders.
Strongly Preferred
Experience with any of the following is highly desirable:
  • Knowledge graphs and ontologies
  • RDF / SPARQL
  • Neo4j
  • Graph modeling over a lakehouse
Enterprise-scale text-to-SQL solutionsSemantic-layer-backed natural-language query platformsML-based anomaly detection for:Time-series dataFinancial transactionsFinancial services or insurance domain experience, particularly:Finance closeGeneral LedgerSubledgerReconciliationActuarial dataLLMOps / MLOps, including:Model versioningPrompt versioningCost governanceObservabilityCertifications such as:Databricks Data Engineer ProfessionalMicrosoft Azure DP-203Microsoft Fabric DP-700AZ-305Experience with:dbtGreat Expectations or similar data-quality platformsWorkdayWorkday PrismWorkday Accounting Center Submission Requirements
Please include the following information:
  • Updated resume
  • Full Name
  • Current Location
  • Contact Number
  • Email Address
  • Work Authorization: U.S. Citizen
  • LinkedIn Profile
  • Availability to Start
  • Interview Availability for the next 3 days
We are looking for a true hands-on Lead Data Engineer who can combine enterprise architecture ownership with day-to-day production engineering. Candidates should be equally comfortable designing the target architecture, reviewing technical decisions, troubleshooting production workloads, and writing production code.

Resumen del puesto

Tipo de empleo

Tiempo completo

Correo electrónico

business@ribbitzllc.com

Habilidades requeridas

PythonScalaPySparkSpark workload troubleshooting and performance tuningDelta Lake and lakehouse architecture (Bronze/Silver/Gold, partitioning, schema evolution, retention)Databricks (Jobs & Workflows, Model Serving, Asset Bundles, cluster and pool management)ADLS Gen2 and Azure data stack (zone architecture, ACLs, lifecycle management)Streaming and ingestion frameworks (Kafka, Azure Event Hubs, CDC, Spark Structured Streaming)SQL and data modeling (dimensional and normalized modeling, schema design, data contracts)CI/CD and Infrastructure as Code (Azure DevOps, Terraform, automated testing for data pipelines)Generative AI / LLM engineering (RAG, embeddings, chunking, LangChain/LlamaIndex/agentic workflows, prompt engineering)AI/ML-based data quality and reconciliation (data drift, schema drift, anomaly detection, synthetic test data)Metadata and data governance (schema inference, profiling, lineage, data catalogs, Unity Catalog)Azure security and identity (Microsoft Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling)Technical leadership and communication (mentoring, code reviews, architecture reviews, presenting to technical and non-technical stakeholders)

Empleos similares

I8IS INC.

Senior Product Owner – ERP Finance Transformation

I8IS INC.

Cary, USPresencialContratoTiempo completo
hace 17 días
I8IS INC.

API Architect – Azure

I8IS INC.

Cary, USPresencialContratoTiempo completo
hace 19 días
OTSI

Senior Product Manager

OTSI

Cary, USPresencialContratoTiempo completo
hace 23 días
Atlantic MEDsearch

Internal Medicine Job Near Cary, NC

Atlantic MEDsearch

Cary, USPresencialContratoTiempo completo
hace 14 días
Atlantic MEDsearch

Family Practice Job Near Cary, NC

Atlantic MEDsearch

Cary, USPresencialContratoTiempo completo
hace 15 días
Simple Solutions

Site Reliability Engineer (SRE) – Security Infrastructure- Hybrid

Simple Solutions

Cary, USPresencialContratoTiempo completo
hace 23 días
Lead Data Engineer (Hands-On) en RibbitZ en Cary | Scovai | Scovai