Senior AI Data Scientist, Agentic Automation
Il y a 9 heures
Ghent, Flanders, Belgique
team.blue
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
Company overview
team.blue is the market leader in enabling digital success for small and medium-sized businesses (SMBs) across Europe, catering to over 3 million customers in 25+ languages. Our mission is to make online business success simpler, by providing our customers with all the tools and resources they need to excel online and remain ahead of the curve.
Position overview We are looking for a senior data scientist to streamline marketing operations at team.blue, by building agentic systems to run them. You would report into the Applied AI team and work on marketing automation projects.
Marketing here runs across many brands, markets and languages, on a stack that differs brand by brand. The work spans competitive and pricing monitoring, performance reporting and diagnosis, SEO and AI-answer visibility, content refresh, localisation and lifecycle production, paid search and social account hygiene, and tracking and consent QA. Each of these is a multi-step process across several systems, repeated per brand.
The method you will follow matters more than the domain: map processes, quantify the time and resources they consume, determine the ROI impact of agentic automation, build a proof of concept, take it to production and measure the impact of your work. This is closer to building autonomous, business-impact systems than to building pure single purpose models.
What we are actually screening for Classical ML and applied statistics are the entry fee for this role, necessary and assumed. Everyone we are talking to has them.
Four things separate candidates.
Can you build an agent someone should trust Most of these agents produce a judgment, backed by numbers: this page lost traffic because of a SERP change, this brand under-converts relative to a comparable one, this competitor’s pricing move matters. A confidently wrong judgment gets read and acted on across several markets before anyone checks it. Lots of this is irreversible, so a recommendation nobody can reconstruct the reasoning for is worse than no recommendation, because it costs trust. Calibration, provenance and auditable workflows are key components of the systems we develop, not a compliance layer on top of them.
Can you tell whether the data underneath is worth reasoning over These agents read from analytics, search console, ad platforms, CRM and third-party SEO and social tools. Tracking is inconsistent across brands, UTM conventions are followed unevenly, and tags break silently. An agent built on that without checking will generate fluent nonsense at scale. Part of the work is refusing to build on a source until it is trustworthy, and saying so with evidence.
Can you reshape a request You will be handed requests written by domain experts, and some of them will be the wrong shape: an agent asked to do something a query would do better, or scoped to advise where it could act. We need someone who can understand that, explain better ways to structure the process, and propose a version that works, rather than building what was asked and shipping a thing nobody uses.
Can you take it to production yourself We mean end to end literally. You write it, you containerise it, you instrument it, and you deploy it with minimal guidance from the devops teams. If the last three things you built were Jupyter notebooks handed to someone else to productionise, this is the wrong role, and no amount of modelling depth compensates.
Your day would involve
Sitting with an SEO or paid search owner and mapping how a traffic-drop investigation actually runs today across brands, then attaching hours per week to each step of it
Extracting requirements live from people who do not think in data models
Designing the state transitions: what triggers, what branches, which APIs get called, where it waits for a human, and what happens when step 4 of 9 fails or a vendor rate-limits you mid-run
Building the guardrails before the capability: dry-run mode, an approval gate ahead of anything that writes to a live account or publishes externally, least-privilege API scopes, a documented undo
Deciding where a human stays in the loop, at what confidence threshold, and designing a review queue marketers will open a second time
Writing evals for output that precision and recall do not capture: is the diagnosis correct, is the cited source real and does it say what the agent claims, does a generated brief hold brand voice in Greek and Dutch as well as in English
Checking whether the tracking data an agent depends on is sound before building on it, and quantifying the error when it is not
Wiring an agent to a webhook or a scheduled trigger, and making the handler idempotent so a retry does not double-post a recommendation or apply the same keyword exclusion twice
Deciding which steps in a flow warrant a frontier model and
Position overview We are looking for a senior data scientist to streamline marketing operations at team.blue, by building agentic systems to run them. You would report into the Applied AI team and work on marketing automation projects.
Marketing here runs across many brands, markets and languages, on a stack that differs brand by brand. The work spans competitive and pricing monitoring, performance reporting and diagnosis, SEO and AI-answer visibility, content refresh, localisation and lifecycle production, paid search and social account hygiene, and tracking and consent QA. Each of these is a multi-step process across several systems, repeated per brand.
The method you will follow matters more than the domain: map processes, quantify the time and resources they consume, determine the ROI impact of agentic automation, build a proof of concept, take it to production and measure the impact of your work. This is closer to building autonomous, business-impact systems than to building pure single purpose models.
What we are actually screening for Classical ML and applied statistics are the entry fee for this role, necessary and assumed. Everyone we are talking to has them.
Four things separate candidates.
Can you build an agent someone should trust Most of these agents produce a judgment, backed by numbers: this page lost traffic because of a SERP change, this brand under-converts relative to a comparable one, this competitor’s pricing move matters. A confidently wrong judgment gets read and acted on across several markets before anyone checks it. Lots of this is irreversible, so a recommendation nobody can reconstruct the reasoning for is worse than no recommendation, because it costs trust. Calibration, provenance and auditable workflows are key components of the systems we develop, not a compliance layer on top of them.
Can you tell whether the data underneath is worth reasoning over These agents read from analytics, search console, ad platforms, CRM and third-party SEO and social tools. Tracking is inconsistent across brands, UTM conventions are followed unevenly, and tags break silently. An agent built on that without checking will generate fluent nonsense at scale. Part of the work is refusing to build on a source until it is trustworthy, and saying so with evidence.
Can you reshape a request You will be handed requests written by domain experts, and some of them will be the wrong shape: an agent asked to do something a query would do better, or scoped to advise where it could act. We need someone who can understand that, explain better ways to structure the process, and propose a version that works, rather than building what was asked and shipping a thing nobody uses.
Can you take it to production yourself We mean end to end literally. You write it, you containerise it, you instrument it, and you deploy it with minimal guidance from the devops teams. If the last three things you built were Jupyter notebooks handed to someone else to productionise, this is the wrong role, and no amount of modelling depth compensates.
Your day would involve
Sitting with an SEO or paid search owner and mapping how a traffic-drop investigation actually runs today across brands, then attaching hours per week to each step of it
Extracting requirements live from people who do not think in data models
Designing the state transitions: what triggers, what branches, which APIs get called, where it waits for a human, and what happens when step 4 of 9 fails or a vendor rate-limits you mid-run
Building the guardrails before the capability: dry-run mode, an approval gate ahead of anything that writes to a live account or publishes externally, least-privilege API scopes, a documented undo
Deciding where a human stays in the loop, at what confidence threshold, and designing a review queue marketers will open a second time
Writing evals for output that precision and recall do not capture: is the diagnosis correct, is the cited source real and does it say what the agent claims, does a generated brief hold brand voice in Greek and Dutch as well as in English
Checking whether the tracking data an agent depends on is sound before building on it, and quantifying the error when it is not
Wiring an agent to a webhook or a scheduled trigger, and making the handler idempotent so a retry does not double-post a recommendation or apply the same keyword exclusion twice
Deciding which steps in a flow warrant a frontier model and