System-Level Reliability Researcher
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Overview
Join imec’s center of excellence for hardware-software-technology co-design to define the future of system-level reliability in compute systems. Compute System Architecture (CSA) is a center of excellence at imec for hardware-software-technology co-design for future compute systems. We work in close collaboration with other expertise centers in imec specializing in applications, technology, circuits and design to innovate and pathfind next-generation compute system architectures across multiple domains – AI, HPC, Automotive, Space, and more. CSA has presence in 6 centers of IMEC, in Belgium, Netherlands, Germany, UK, USA and Qatar. This position is primarily for Leuven, Belgium.
What you will do
We are looking for an experienced System-Level Reliability Researcher to join our team. In this position, you will play a key role in developing methodologies to evaluate and improve the reliability of advanced compute architectures. You will focus on how hardware faults, degradation effects, and technology-level reliability observations translate into system-level behavior, application correctness, availability, and product lifetime.
This R&D role focuses on compute system architecture, reliability modeling, and resilience evaluation. You will collaborate with experts from the technology, device, circuit, and design domains to translate lower-level reliability information into architectural fault models, simulation inputs, and system-level reliability metrics.
Responsibilities
- Developing system-level reliability assessment methods for processors, accelerators, memory subsystems, SoCs, and chiplet-based architectures.
- Creating architectural fault models and abstractions based on reliability information provided by lower layers of the technology stack.
- Building and using simulation, emulation, and fault-injection frameworks to study fault propagation, error masking, and application-level impact.
- Quantifying reliability outcomes such as Silent Data Corruption, detected errors, service interruptions, availability loss, and lifetime-related degradation at the system level.
- Evaluating Reliability, Availability, and Serviceability (RAS) mechanisms, including error detection, containment, recovery, redundancy, and graceful degradation techniques.
- Studying workload-dependent reliability behavior under realistic execution conditions and operating profiles.
- Connecting technology-informed fault characteristics with system architecture models to support reliability-aware design decisions.
- Collaborating closely with colleagues within and outside CSA to interface with workload models, architectural simulators, circuit-level reliability data, and technology-level observations.
- Contributing to research publications, partner discussions, etc.