AgileEngine
Middle Data Engineer ID86295
Bogota, COOp locatieVastFulltime
Geplaatst 29 aug 2026
Deze baan is geplaatst in het ES
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture across bronze, silver, and gold layers. This role also uses AI tools like Claude and GitHub Copilot to speed up development.
WHAT YOU WILL DO
- Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
- Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
- Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
- Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
- Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
- Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
- Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
- Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
- Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.
MUST HAVES
- 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
- Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
- Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
- Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
- Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
- Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
- Strong problem-solving, collaboration, and communication skills.
- Familiarity with Unity Catalog, data governance, access control, and PII handling.
- Experience with dbt or an equivalent transformation framework.
- Familiarity with secure coding standards and industry security best practices.
- Experience delivering production data platforms at scale.
- Upper-intermediate English level.
NICE TO HAVES
- Experience with Infrastructure as Code (IaC) using Terraform and CI/CD using Azure DevOps.
- Experience working with relational databases (specifically PostgreSQL) and data persistence concepts.
- Familiarity with logging and monitoring tools (e.g., Dynatrace, CloudWatch, Databricks system tables).
- Experience working in Agile or team-based development environments preferred.
PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture across bronze, silver, and gold layers. This role also uses AI tools like Claude and GitHub Copilot to speed up development.
WHAT YOU WILL DO
- Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
- Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
- Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
- Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
- Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
- Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
- Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
- Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
- Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.
MUST HAVES
- 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
- Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
- Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
- Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
- Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
- Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
- Strong problem-solving, collaboration, and communication skills.
- Familiarity with Unity Catalog, data governance, access control, and PII handling.
- Experience with dbt or an equivalent transformation framework.
- Familiarity with secure coding standards and industry security best practices.
- Experience delivering production data platforms at scale.
- Upper-intermediate English level.
NICE TO HAVES
- Experience with Infrastructure as Code (IaC) using Terraform and CI/CD using Azure DevOps.
- Experience working with relational databases (specifically PostgreSQL) and data persistence concepts.
- Familiarity with logging and monitoring tools (e.g., Dynatrace, CloudWatch, Databricks system tables).
- Experience working in Agile or team-based development environments preferred.
PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location
Rolschets
Type baan
Fulltime
Vereiste vaardigheden
Apache Spark (PySpark)Databricks platform (including Databricks Workflows)Delta LakePython programmingAdvanced SQLData modeling (dimensional and Lakehouse/medallion architecture)Streaming ingestion (Structured Streaming, Auto Loader, Kafka, Event Hubs)Workflow orchestration (Airflow, Azure Data Factory, Databricks Workflows)Data governance and access control (Unity Catalog, PII handling)Data quality, lineage, and automated validation checksSpark and Delta performance optimization (partitioning, clustering, caching, cluster sizing)Debugging and troubleshooting production data pipelines and incidentsdbt or equivalent transformation frameworksLegacy ETL/data warehouse migration and modernizationUse of AI-assisted development tools (Claude, GitHub Copilot)
Vergelijkbare banen
Senior Data Engineer ID86297
AgileEngine
Bogota, COOp locatieVastFulltime
21 uur geleden
Full Stack Software Engineer ID85839
AgileEngine
Bogota, COOp locatieVastFulltime
21 uur geleden
Senior Quality Engineer ID84192
AgileEngine
Bogota, COOp locatieVastFulltime
gisteren
Senior Data Scientist – MMM ID71005
AgileEngine
Bogota, COOp locatieVastFulltime
gisteren
Service Appointment Coordinator ID85149
AgileEngine
Bogota, COOp locatieVastFulltime
eergisteren
Product Owner ID85383
AgileEngine
Bogota, COOp locatieVastFulltime
eergisteren