System-Level Reliability Researcher

Il y a 3 jours

Heverlee, Vlaams-Brabant, Belgique Leuven Temps plein 90 000 € - 120 000 € Contrat

Overview

Join imec’s center of excellence for hardware-software-technology co-design to define the future of system-level reliability in compute systems. Compute System Architecture (CSA) is a center of excellence at imec for hardware-software-technology co-design for future compute systems. We work in close collaboration with other expertise centers in imec specializing in applications, technology, circuits and design to innovate and pathfind next-generation compute system architectures across multiple domains – AI, HPC, Automotive, Space, and more. CSA has presence in 6 centers of IMEC, in Belgium, Netherlands, Germany, UK, USA and Qatar. This position is primarily for Leuven, Belgium.

What you will do

We are looking for an experienced System-Level Reliability Researcher to join our team. In this position, you will play a key role in developing methodologies to evaluate and improve the reliability of advanced compute architectures. You will focus on how hardware faults, degradation effects, and technology-level reliability observations translate into system-level behavior, application correctness, availability, and product lifetime.

This R&D role focuses on compute system architecture, reliability modeling, and resilience evaluation. You will collaborate with experts from the technology, device, circuit, and design domains to translate lower-level reliability information into architectural fault models, simulation inputs, and system-level reliability metrics.

Responsibilities

  • Developing system-level reliability assessment methods for processors, accelerators, memory subsystems, SoCs, and chiplet-based architectures.
  • Creating architectural fault models and abstractions based on reliability information provided by lower layers of the technology stack.
  • Building and using simulation, emulation, and fault-injection frameworks to study fault propagation, error masking, and application-level impact.
  • Quantifying reliability outcomes such as Silent Data Corruption, detected errors, service interruptions, availability loss, and lifetime-related degradation at the system level.
  • Evaluating Reliability, Availability, and Serviceability (RAS) mechanisms, including error detection, containment, recovery, redundancy, and graceful degradation techniques.
  • Studying workload-dependent reliability behavior under realistic execution conditions and operating profiles.
  • Connecting technology-informed fault characteristics with system architecture models to support reliability-aware design decisions.
  • Collaborating closely with colleagues within and outside CSA to interface with workload models, architectural simulators, circuit-level reliability data, and technology-level observations.
  • Contributing to research publications, partner discussions, etc.