Senior Data Engineer

Il y a 2 jours

, Belgique CluePoints Temps plein 75 000 € - 110 000 € Contrat

Job Description

At CluePoints, we're redefining how clinical trials are run. As a leading provider of Risk-Based Quality Management and Data Quality Oversight software, we use advanced statistics, artificial intelligence and machine learning to improve the quality, accuracy and integrity of clinical trial data - turning artificialintelligence into human intelligence. We're an ambitious, fast-growing technology company with a dynamic and diverse international team. Collaboration, flexibility and continuous learning are part of our DNA. Guided by our values of Care, Passion and Smart Disruption, we're united by a shared mission: to create smarter ways to run efficient clinical trials and deliver AI-powered insights that improve human outcomes worldwide.

Job Description

At CluePoints, we're redefining how clinical trials are run. As a leading provider of Risk-Based Quality Management and Data Quality Oversight software, we use advanced statistics, artificial intelligence and machine learning to improve the quality, accuracy and integrity of clinical trial data - turning artificialintelligence into human intelligence. We're an ambitious, fast-growing technology company with a dynamic and diverse international team. Collaboration, flexibility and continuous learning are part of our DNA. Guided by our values of Care, Passion and Smart Disruption, we're united by a shared mission: to create smarter ways to run efficient clinical trials and deliver AI-powered insights that improve human outcomes worldwide.

The Role

We're looking for a hands-on Senior Data Engineer to own the technical design, reliability and evolution of our Databricks data foundation. You'll work alongside Clinical Data Insights Analysts - colleagues with strong business, clinical and BI expertise who also write SQL and Python - complementing their skills with deeper engineering, automation and production-grade rigour.

Job Requirements

  • 5+ years in data engineering or data-platform engineering.
  • Strong hands-on Databricks experience (or comparable cloud lakehouse platform).
  • Advanced SQL and strong Python / PySpark skills.
  • Production-grade pipeline experience - design, build, test, deploy and maintain.
  • Solid understanding of Delta Lake, medallion architecture, and OLAP/OLTP transformation.
  • Experience with data governance, Unity Catalog (or equivalent), automated quality controls and CI/CD.
  • At least one major cloud platform: Azure, AWS or GCP.
  • Practical use of generative AI in engineering workflows, with solid critical judgement on its outputs.
  • High autonomy - you take ownership from design through deployment, monitoring and maintenance.
  • Collaborative by nature; comfortable working directly alongside analysts who also write code.

Nice to Have

  • Lakeflow Jobs / Lakeflow Declarative Pipelines; Databricks Asset Bundles.
  • Databricks AI/BI, Genie, conversational analytics or AI agent experience.
  • Databricks Data Engineer certification.
  • Experience in a regulated environment (clinical trials, life sciences or similar).

Tech Stack

  • Delta Lake
  • Unity Catalog
  • Lakeflow Jobs & Declarative Pipelines
  • Medallion architecture
  • SQL / Python / PySpark
  • Microsoft Azure
  • Git / CI/CD
  • Databricks Asset Bundles
  • Databricks AI/BI
  • Genie

Job Responsibilities

  • Own the architecture, reliability and governance of the team's Databricks lakehouse (Delta Lake, medallion architecture, Unity Catalog).
  • Design and operate batch and streaming ingestion pipelines from the CluePoints platform and other business systems (e.g. Zendesk).
  • Co-develop SQL, Python and PySpark transformations with the analyst team; review and optimize for performance, scalability and cost.
  • Productionize analytical and AI prototypes - adding testing, orchestration, monitoring, error handling and deployment controls.
  • Build automated data-quality controls, monitoring and alerting; investigate and resolve pipeline incidents.
  • Implement technical governance: naming conventions, metadata, lineage, tagging and access controls via Unity Catalog.
  • Apply AI and automation across the data lifecycle - code development, documentation, metadata classification and inconsistency detection.
  • Collaborate with Engineering on source-system access and with Product when Databricks insights are candidates for platform integration.

Job Benefits

What We Offer - Belgium

  • Health Insurance through Alan (100% hospitalisation cover, 80% ambulatory and dental)
  • Mobility Budget for eco-transport, housing, or car allowance (flexible 3-pillar system)
  • Group Insurance Plan with 6-12% employer pension contribut