Jobtailor
Portugal / Global
Portugal / Global
Build and maintain scalable data processing pipelines and workflows on the Databricks Lakehouse Platform using Apache Spark and PySpark
Ingest and integrate datasets from databases, file systems, and APIs into the Lakehouse using Auto Loader, CDC, and batch or streaming ingestion patterns
Model and implement bronze, silver, and gold transformation layers using Delta Lake and Databricks SQL
Ensure data integrity, consistency, and quality through validation and monitoring using Delta Live Tables expectations and Lakehouse monitoring capabilities
Tune pipelines for performance and cost through Spark job optimization, Delta table maintenance, and cluster and compute configuration
Apply data security, access control, and privacy practices using Unity Catalog for governance, permissions, and lineage
Contribute to CI/CD and deployment automation for data assets
Support migration initiatives from on-premise Cloudera environments to the Databricks Lakehouse
Collaborate with data scientists and analysts to deliver reliable datasets for analytics and machine learning workflows
Occasionally support lightweight Python services and APIs exposing data to downstream consumers
Requirements
Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field
2+ years of professional experience in data engineering or software engineering with a strong data component
Solid Python skills in production environments
Good engineering practices, including version control, testing, and code review
Hands-on experience with Spark/PySpark for large-scale batch or streaming data processing
Working experience with Databricks or a comparable Lakehouse/cloud data platform
Experience with one major cloud provider: Azure, AWS, or GCP
Strong SQL skills
Solid data modelling fundamentals across relational and non-relational stores
Fluency in English, written and spoken
Databricks certifications are nice to have
Experience with Delta Live Tables, Structured Streaming, or Unity Catalog is nice to have
Familiarity with Databricks Workflows, Airflow, Kafka, Terraform, Databricks Asset Bundles, CI/CD tooling, Photon, Databricks SQL Warehouses, or serverless compute is nice to have
Core Competencies
Demonstrates expertise in building and maintaining scalable data processing pipelines on the Databricks Lakehouse Platform, utilizing Apache Spark and PySpark. Proficient in data ingestion, transformation, and ensuring data integrity while applying best practices in data security and governance.
Highest-signal resume keywords
Apache Spark
PySpark
Databricks Lakehouse
SQL
Data Engineering
Hard Skills
Python
Data Modeling
Delta Lake
Delta Live Tables
Structured Streaming
CI/CD
Version Control
Testing
Code Review
Data Ingestion
Soft Skills
Collaboration
Communication
Certifications & Qualifications
Databricks Certifications
Industry Keywords
Cloud Data Platform
Data Processing Pipelines
Data Integrity
Data Quality
Data Security
Tools & Technologies
Auto Loader
Unity Catalog
Airflow
Kafka
Terraform
#J-18808-Ljbffr
Lisboa / Global
Lisboa / Global
Portugal / Global
Lisboa / Global
Portugal / Global
Portugal / Global