AdvanceWorks
Lisboa / Global
Databricks Engineer
- Hybrid
Lisboa / Global
Challenge yourself and us At AdvanceWorks, we learn together, grow together, and have fun together. And we do all this while creating great software solutions for clients across many industries and geographies.
We are constantly challenging ourselves and each other to use the latest technologies and the best methodologies and make sure we walk the talk.
What will be your role
We are looking for a Databricks Data Engineer to join our team in Lisbon and contribute to the development and evolution of modern data ingestion and Change Data Capture (CDC) solutions for a global organization operating at scale.
In this role, you'll work on a modern data architecture built around Databricks, Apache Spark, Delta Lake, Azure, and Kafka, helping evolve a flexible and scalable approach to data ingestion across multiple source systems.
You'll have an active role in defining and evolving ingestion patterns, developing Databricks notebooks and processing frameworks, configuring pipelines, and ensuring the integrity, traceability, and scalability of data throughout the process.
The project is moving from a traditional CDC approach focused primarily on Mainframe environments towards a more flexible architecture capable of supporting multiple data sources and destinations. This creates an opportunity to work on a technically challenging environment where Big Data, Cloud, CDC, and modern data engineering practices come together.
Your responsibilities will include
Developing and maintaining Change Data Capture (CDC) applications and data ingestion processes in Databricks
Creating and maintaining Databricks notebooks for setup, configuration, and data processing
Developing and parametrizing batch processing pipelines using Apache Spark
Implementing CDC logic to correctly detect and process inserts, updates, and deletes
Ensuring the generation of complete data snapshots while maintaining historical changes and data traceability
Working with Delta Lake to support reliable, scalable, and performant data processing
Working with multiple data sources, including Azure File Share, DB2, SQL Server, and SingleStore
Integrating processed data with destinations such as SingleStore and Kafka Confluent
Working with Databricks Workflows and Azure services to orchestrate and operationalize data pipelines
Optimizing processing performance and ensuring that pipelines are scalable and resilient
Contributing to the definition and evolution of data ingestion architecture and processing frameworks
Collaborating with other Data Engineers and technical teams to integrate processed data with downstream microservices
Ensuring data integrity, consistency, and traceability throughout migration and processing flows
Applying best practices around ETL/ELT, data engineering, monitoring, error handling, and pipeline reliability
Contributing to continuous improvement initiatives across the data platform and its ingestion frameworks
What should you bring to the team
3+ years of experience in Data Engineering, Databricks, or related data-intensive roles
Solid hands-on experience with Databricks, including notebooks, clusters, jobs, and workflows
Strong knowledge of Apache Spark, including batch processing and distributed data transformations
Strong experience with Python and SQL
Hands-on experience with Delta Lake and modern lakehouse architectures
Practical understanding of Change Data Capture (CDC) concepts and implementation patterns
Experience developing and parametrizing data pipelines and processing frameworks
Experience working with cloud data platforms, particularly Microsoft Azure
Familiarity with Azure services such as Azure File Share and Data Lake
Experience integrating data pipelines with Kafka, ideally Kafka Confluent
Good understanding of data architecture and ETL/ELT best practices
Experience designing scalable and resilient data processing solutions
Strong problem-solving skills and ability to troubleshoot distributed data processing environments
Strong communication skills and ability to collaborate with different technical teams
Availability to work in a hybrid model, 3 days per week in Tagus Park
Fluent Portuguese and English, both written and spoken
Nice to have
Experience with SingleStore
Experience with SQL-based CDC implementations across different database technologies
Experience with file-based CDC or ingestion frameworks
Knowledge of microservices and event-driven architectures
Experience with MLOps or Data Engineering in cloud environments
Familiarity with AI/GenAI pipelines and data platforms supporting AI initiatives
Experience with streaming technologies and event-driven data architectures
Knowledge of performance optimization techniques for Spark and Databricks
Experience designing resilient and highly scalable data solutions
What is in it for you
An amazing informal culture of smart, hardworking, and friendly people who support and care about each other
A mentorship program from day one
Opportunity to work on an innovative project combining CDC, Big Data, Cloud, Databricks, and Kafka
The chance to contribute to the evolution of a modern and scalable data architecture
Formal training and certifications in data, cloud, and engineering excellence
Access to cutting-edge tools and technologies
A dynamic company where your ideas matter more than job titles
Flexibility with responsibility
A Rubber Duck to help you debug those tricky Spark transformations
Don't hold off any longer and apply now!
If you have any questions, drop us a line at [email protected]
Portugal / Global
Portugal / Global
Lisbon / Global
Lisboa / Global
Portugal / Global
Portugal / Global