Data & AI Engineer

Il y a 2 jours

Brussel, Brussel-Hoofdstad, Belgique Appsierra Group Temps plein 65 000 € - 90 000 € Contrat

Language: English

2 days a week onsite

Only EU Citizen

Tasks:

  • Develop and maintain data pipelines that integrate multiple heterogeneous data sources, including both structured and unstructured information.
  • Implement data ingestion processes, including batch and near-real-time processing.
  • Perform data cleansing, validation and standardization to ensure reliable and consistent datasets.
  • Apply metadata tagging and support data lineage across relevant data flows.
  • Contribute to the development and maintenance of the Data and Product Catalogue.
  • Implement data quality checks, validation rules and monitoring mechanisms across data pipelines and analytical workflows.
  • Support the identification and remediation of data quality issues in cooperation with Data Stewards and relevant governance teams.
  • Contribute to the creation of reusable datasets and data products from heterogeneous information sources.
  • Design solutions that combine structured data, such as databases and tabular datasets, with unstructured data, such as documents, reports and text.
  • Transform unstructured information into formats suitable for analysis, reporting and AI enabled processing.
  • Enable unified analytical workflows and reporting across mixed data types.
  • Support AI-driven processing of document-centric data.
  • Implement analytics capabilities that support operational and strategic decision-making workflows.
  • Contribute to the design and implementation of AI-enabled use cases.
  • Ensure AI outputs are explainable, traceable and supported by appropriate human-in the-loop controls.
  • Automate data pipelines, reporting workflows and recurring analytical processes.
  • Implement event-driven processing, alerts and triggers where relevant.
  • Support monitoring, logging and operational observability of platform processes.
  • Implement security controls aligned with EU requirements, including identity and access management, encryption of data at rest and in transit, and audit logging.
  • Support the separation of classified and unclassified environments.
  • Contribute to solutions that can be deployed in secure or air-gapped environments.
  • Build modular and scalable platform components using open standards and APIs.
  • Contribute to integration with existing EDA systems and external data sources.
  • Support hybrid and sovereign deployment approaches.
  • Produce clear technical documentation and support knowledge transfer to relevant stakeholders.

Mandatory Requirements:

Data Engineering & Architecture

  • Proven experience in designing and implementing data pipelines, including ETL/ELT.
  • Strong knowledge of data lake and data warehouse architectures.
  • Experience with modern data platforms, such as Microsoft Fabric, Copilot, the Azure ecosystem, and open-source data platforms.

Handling Structured & Unstructured Data (Critical Requirement)

  • Demonstrated experience in handling and integrating structured data, such as databases and tabular datasets, and unstructured data, such as documents, reports, PDFs, and text corpora.
  • Ability to build pipelines enabling end-to-end exploitation of heterogeneous data.
  • Experience preparing data for reporting and advanced analytics.
  • Ability to structure unstructured information using metadata, classification, and transformation techniques.

Analytics & AI

  • Experience with analytics development and data modelling.
  • Exposure to AI/ML solutions, particularly on text- or document-based data.
  • Understanding of explainability and traceability principles.

Engineering Practices

  • Knowledge of DevOps, Git, CI/CD, and pipeline automation.
  • Ability to deliver in multi-stakeholder environments.

Additional Desirable Skills:

  • Experience with classified or restricted environments, such as EUCI or equivalent.
  • Familiarity with Microsoft Purview or similar governance tools.
  • Ability to create and work with MCP servers.
  • Experience with data catalogues and business glossaries.
  • Experience with cross-domain data handling.
  • Experience with hybrid or sovereign cloud environments.
  • Experience in public sector, or EU institutions.
  • Experience in AI explainability or human-in-the-loop systems.