Senior Data Scientist, AI Scoring
Il y a 3 heures
Brussels, Brussels, Belgique
Workera AI
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
Senior Data Scientist - Assessment Scoring & Evaluation
Everyone’s racing to build AI. Workera exists for the 8 billion people who work alongside it.
While the world’s attention is on creating new tools, someone has to solve the other side of the equation: the humans. The workforce is going through the biggest transformation in a generation and most organizations are navigating it blind, without the data to understand what their people can actually do, where the gaps are, or how to close them fast enough.
As Workera moves into higher-stakes decisions, you’re the person who can prove a measurement is accurate, fair, and defensible. You’ll own the measurement system that turns signals into the skill data our customers act on: the evaluation harnesses, the calibration methods, the quality monitoring behind every score. You’ll also shape how our AI software itself is designed, from how components are orchestrated and prompted to the evaluation frameworks that decide whether they’re good enough to ship.
You’re not just running analyses, you’re building the trust layer underneath a product that increasingly makes high-stakes calls about people’s careers. If you want a role where your judgment becomes the thing customers, auditors, and your own engineering team rely on, and where the measurement problems are still being invented rather than optimized, this is it.
Workera’s skills intelligence platform is critical infrastructure for the AI era: the layer that lets organizations understand, mobilize, manage, and develop their talent with precision. We’re trusted by the Fortune 500, powered by proprietary AI agents, and built by a small, senior team, which means what you ship here has outsized reach.
WHY THIS ROLE EXISTS Workera is scaling fast: more customers, more use cases, more fields evaluated, higher stakes on every measurement. That scale changes what quality means. What we once crafted and inspected by hand now needs monitoring that catches issues before customers do, and remediation that resolves them fast.
This role sits at the intersection of Engineering, Assessment Science, and Design, partnering closely with the product engineers who build our scoring pipeline. You’re the data layer that pieces those disciplines together: the foundation the rest of our scoring is built on, and the reason a construct definition from Assessment Science actually turns into a number a customer can trust.
YOUR TEAM You report to Dr. Taylor Sullivan, Workera’s VP of Product and Assessments, and work day to day with two groups: The assessment science team decides what we’re measuring and how a skill turns into test content, they own construct definition (what “good at X” actually means), blueprinting (how an assessment is structured), and content authoring (writing the actual questions). You partner with them on how scoring methods serve that intent. The assessment tech team builds and maintains the scoring pipeline; you work embedded with them: writing specs, implementing and coordinating improvements, and validating that what ships meets the quality bar. You also work with data engineering on data availability. It’s a small, cross-functional team, and you sit inside product decisions rather than in a separate research function.
WHAT YOU’LL OWN Build trust by owning our scoring mechanism. These are the outcomes you’re accountable for:
Own scoring quality end to end: evaluator design, rubric anchoring, calibration, and accountability for accuracy against expert benchmarks
Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection, and the pre-release gate that every scoring change must clear
Define and publish assessment quality KPIs (human to AI agreement, reliability, classification accuracy, bias indicators, latency, cost per assessment) on dashboards anyone in the company can reference
Lead the analytics investigations behind assessment decisions: performance studies, impact simulations, root cause analysis when scores look wrong
Ship measurement improvements end to end: implement and coordinate with tech, own the spec, the validation, and the quality bar in both cases
Be Workera’s liaison between AI governance and InfoSec: stay current on frameworks like GDPR and the EU AI Act, and turn that into the evidence we show in audits, security reviews, and enterprise diligence when customers ask how we use AI in our assessments
Empower partners across Product, Engineering, and GTM to get the data and analysis they need autonomously, by building AI tooling and documentation rather than answering each request yourself
HOW YOU’LL RAMP We don’t expect you to figure it out alone. Here’s what great looks like at each stage:
First 30 Days: Learn the Machine
Immerse yourself in Workera’s platform, customers, and the problems we’re solving. You’
While the world’s attention is on creating new tools, someone has to solve the other side of the equation: the humans. The workforce is going through the biggest transformation in a generation and most organizations are navigating it blind, without the data to understand what their people can actually do, where the gaps are, or how to close them fast enough.
As Workera moves into higher-stakes decisions, you’re the person who can prove a measurement is accurate, fair, and defensible. You’ll own the measurement system that turns signals into the skill data our customers act on: the evaluation harnesses, the calibration methods, the quality monitoring behind every score. You’ll also shape how our AI software itself is designed, from how components are orchestrated and prompted to the evaluation frameworks that decide whether they’re good enough to ship.
You’re not just running analyses, you’re building the trust layer underneath a product that increasingly makes high-stakes calls about people’s careers. If you want a role where your judgment becomes the thing customers, auditors, and your own engineering team rely on, and where the measurement problems are still being invented rather than optimized, this is it.
Workera’s skills intelligence platform is critical infrastructure for the AI era: the layer that lets organizations understand, mobilize, manage, and develop their talent with precision. We’re trusted by the Fortune 500, powered by proprietary AI agents, and built by a small, senior team, which means what you ship here has outsized reach.
WHY THIS ROLE EXISTS Workera is scaling fast: more customers, more use cases, more fields evaluated, higher stakes on every measurement. That scale changes what quality means. What we once crafted and inspected by hand now needs monitoring that catches issues before customers do, and remediation that resolves them fast.
This role sits at the intersection of Engineering, Assessment Science, and Design, partnering closely with the product engineers who build our scoring pipeline. You’re the data layer that pieces those disciplines together: the foundation the rest of our scoring is built on, and the reason a construct definition from Assessment Science actually turns into a number a customer can trust.
YOUR TEAM You report to Dr. Taylor Sullivan, Workera’s VP of Product and Assessments, and work day to day with two groups: The assessment science team decides what we’re measuring and how a skill turns into test content, they own construct definition (what “good at X” actually means), blueprinting (how an assessment is structured), and content authoring (writing the actual questions). You partner with them on how scoring methods serve that intent. The assessment tech team builds and maintains the scoring pipeline; you work embedded with them: writing specs, implementing and coordinating improvements, and validating that what ships meets the quality bar. You also work with data engineering on data availability. It’s a small, cross-functional team, and you sit inside product decisions rather than in a separate research function.
WHAT YOU’LL OWN Build trust by owning our scoring mechanism. These are the outcomes you’re accountable for:
Own scoring quality end to end: evaluator design, rubric anchoring, calibration, and accountability for accuracy against expert benchmarks
Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection, and the pre-release gate that every scoring change must clear
Define and publish assessment quality KPIs (human to AI agreement, reliability, classification accuracy, bias indicators, latency, cost per assessment) on dashboards anyone in the company can reference
Lead the analytics investigations behind assessment decisions: performance studies, impact simulations, root cause analysis when scores look wrong
Ship measurement improvements end to end: implement and coordinate with tech, own the spec, the validation, and the quality bar in both cases
Be Workera’s liaison between AI governance and InfoSec: stay current on frameworks like GDPR and the EU AI Act, and turn that into the evidence we show in audits, security reviews, and enterprise diligence when customers ask how we use AI in our assessments
Empower partners across Product, Engineering, and GTM to get the data and analysis they need autonomously, by building AI tooling and documentation rather than answering each request yourself
HOW YOU’LL RAMP We don’t expect you to figure it out alone. Here’s what great looks like at each stage:
First 30 Days: Learn the Machine
Immerse yourself in Workera’s platform, customers, and the problems we’re solving. You’