Skip to main content
futureproof logo

Data Engineer (AWS, DBT, Trino/Databricks)

futureproof
2 hours ago
Contract
Remote
Worldwide
Marketing

This is a remote position.

We are looking for an experienced Data Engineer to support the migration and integration of data from multiple enterprise and laboratory systems. The project is primarily focused on dbt, advanced SQL, and data modelling rather than Databricks or large-scale PySpark processing.

The successful candidate will design new data models, develop complex transformations, and consolidate data from several source databases into a unified AWS-based data platform.

Key Responsibilities

  • Design logical and physical data models based on technical and business requirements.
  • Develop, test, and maintain complex data transformations using dbt and SQL.
  • Migrate and consolidate data from multiple databases and source systems.
  • Use Trino/Presto to query, join, and transform data distributed across different sources.
  • Integrate data from enterprise and laboratory systems such as SAP, GLIMS, and Veeva.
  • Build reusable, maintainable, and well-documented transformation models.
  • Validate migrated data and investigate data-quality or consistency issues.
  • Optimize complex SQL queries and transformation processes.
  • Support data lineage, traceability, integrity, and documentation.
  • Work with data stored or processed within AWS, particularly Amazon S3 and Athena.
  • Participate in Agile delivery, code reviews, testing, and CI/CD activities.
  • Collaborate with data architects, analysts, engineers, and pharmaceutical business stakeholders.
  • Develop new solutions rather than only maintaining existing pipelines.


Requirements

  • Strong hands-on experience with dbt.
  • Advanced proficiency in SQL, including:
    • Complex joins and transformations
    • CTEs and window functions
    • Query optimization
    • Data reconciliation and validation
  • Practical experience with data modelling, including the ability to design a model from business or technical requirements.
  • Experience migrating and consolidating data from multiple databases.
  • Experience with distributed query engines such as:
    • Trino
    • Presto
  • Working knowledge of AWS data services, particularly:
    • Amazon S3
    • Amazon Athena
    • AWS Glue
  • Experience building reliable, production-ready data transformation pipelines.
  • Understanding of relational databases and data warehousing concepts.

Additional Relevant Skills

  • Python, Scala, or PySpark for scripting, automation, or supplementary data transformation.
  • Experience with PostgreSQL or another relational target database.
  • Familiarity with Git and CI/CD deployment pipelines.
  • Experience working in an Agile delivery environment.
  • Knowledge of data quality, lineage, governance, and integrity principles.

Domain Experience

Experience integrating data from any of the following is valuable:

  • SAP ERP
  • GLIMS or other LIMS platforms
  • Veeva
  • Pharmaceutical manufacturing or laboratory systems

Experience with GxP regulations and pharmaceutical data-integrity requirements is preferred but not necessarily essential.

Qualifications

Candidates should meet one of the following:

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, with at least 4 years of relevant technical experience; or
  • At least 8 years of equivalent experience in data engineering, data modelling, data integration, or cloud analytics without a degree.

Ideal Candidate

The ideal candidate is a SQL-focused Data Engineer with strong practical experience in dbt and data modelling. They should be comfortable combining data from several databases, designing new data structures, and implementing complex transformations with Trino/Presto in an AWS environment.



Benefits

  • Location: European Union
  • Contract Type: Freelance / Contract
  • Start date: Summer, 2026
  • Time Allocation: 40 hours/week
  • Global Pharmaceutical Company in Prague​