Start Your Search Here

Job Search

Jobtailor

Portugal / Global

Data Engineer, Databricks

Job Description

Build and maintain scalable data processing pipelines and workflows on the Databricks Lakehouse Platform using Apache Spark and PySpark

Ingest and integrate datasets from databases, file systems, and APIs into the Lakehouse using Auto Loader, CDC, and batch or streaming ingestion patterns

Model and implement bronze, silver, and gold transformation layers using Delta Lake and Databricks SQL

Ensure data integrity, consistency, and quality through validation and monitoring using Delta Live Tables expectations and Lakehouse monitoring capabilities

Tune pipelines for performance and cost through Spark job optimization, Delta table maintenance, and cluster and compute configuration

Apply data security, access control, and privacy practices using Unity Catalog for governance, permissions, and lineage

Contribute to CI/CD and deployment automation for data assets

Support migration initiatives from on-premise Cloudera environments to the Databricks Lakehouse

Collaborate with data scientists and analysts to deliver reliable datasets for analytics and machine learning workflows

Occasionally support lightweight Python services and APIs exposing data to downstream consumers

Requirements

Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field

2+ years of professional experience in data engineering or software engineering with a strong data component

Solid Python skills in production environments

Good engineering practices, including version control, testing, and code review

Hands-on experience with Spark/PySpark for large-scale batch or streaming data processing

Working experience with Databricks or a comparable Lakehouse/cloud data platform

Experience with one major cloud provider: Azure, AWS, or GCP

Strong SQL skills

Solid data modelling fundamentals across relational and non-relational stores

Fluency in English, written and spoken

Databricks certifications are nice to have

Experience with Delta Live Tables, Structured Streaming, or Unity Catalog is nice to have

Familiarity with Databricks Workflows, Airflow, Kafka, Terraform, Databricks Asset Bundles, CI/CD tooling, Photon, Databricks SQL Warehouses, or serverless compute is nice to have

Core Competencies

Demonstrates expertise in building and maintaining scalable data processing pipelines on the Databricks Lakehouse Platform, utilizing Apache Spark and PySpark. Proficient in data ingestion, transformation, and ensuring data integrity while applying best practices in data security and governance.

Highest-signal resume keywords

Apache Spark

PySpark

Databricks Lakehouse

SQL

Data Engineering

Hard Skills

Python

Data Modeling

Delta Lake

Delta Live Tables

Structured Streaming

CI/CD

Version Control

Testing

Code Review

Data Ingestion

Soft Skills

Collaboration

Communication

Certifications & Qualifications

Databricks Certifications

Industry Keywords

Cloud Data Platform

Data Processing Pipelines

Data Integrity

Data Quality

Data Security

Tools & Technologies

Auto Loader

Unity Catalog

Airflow

Kafka

Terraform

#J-18808-Ljbffr

Candidatar-se Now

Similar Opportunities

View all jobs