Operations Manager – Digital Platform
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Operations Manager – Digital Platform & DevOps Operations
Location Brussels-50% onsite
Contract 6-12 months contract (potential extensions)
On-call Yes (24/7 escalation rota)
About the role
We are looking for an Operations Manager to bring process, control and an operations mindset to our DevOps and digital-platform operations team. The team runs the container platform, middleware and streaming services, and the operational DevOps of the mission-critical applications hosted on them. The services are operational 24/7.
The team is technically strong and already has a DevOps Lead who provides technical leadership and keeps the team together day to day. That does not change. This role sits between management and the DevOps Lead and owns the operations layer: a single, agreed set of priorities;
Scroll down to find an indepth overview of this job, and what is expected of candidates Make an application by clicking on the Apply button.
defined operational processes;
documentation and procedures at a standard that survives staff turnover;
monitoring and alertingthat tell the truth;
and reporting that gives management a reliable picture of platform health and risk. The goal is to move the team from reactive fire-fighting to planned, controlled operations — without losing the engineering strength and cohesion it already has.
You decide what gets done first and how it is run;
the DevOps Lead decideshow it is engineered and keeps the team moving. Making that split work is the core of the job.
This is a role for a process-minded operations manager with enough technical depth to be credible with senior engineers, and enough maturity to lead through an existing lead rather than around them.
The environment
- Enterprise container platform (Red Hat OpenShift / Kubernetes) on public cloud (Azure), hosting business-critical applications
- Middleware stack: application servers, API gateways, identity and access management, secrets management, messaging
- Event streaming (Apache Kafka) and integration services
- CI/CD and GitOps tooling;
infrastructure-as-code - Observability stack: metrics, logs, traces, APM, dashboards and alerting
- 24/7 operations with on-call, formal incident management and service-level commitments
Key responsibilities
Own the operating model and the priorities
- Line-manage the DevOps Lead: set objectives, coach and support them;
through them, ensure the team has clear roles, workload and on-call coverage - Run a single intake and prioritisation process for all platform work — operational stability, platform improvements and requests from application teams — aligned to organisational priorities rather than individual preference
- Shield the team from unplanned demand, and say no, with reasons, when necessary
- Establish the division of responsibilities between operations management and technical leadership, and keep it clear
Bring the process
- Define, implement and enforce operational processes for the platform: incident, problem, change and release management, on-call handover, capacity and patch management — with clear roles and decision rights
- Set the standard for operational documentation: runbooks, standard operating procedures, architecture and configuration records, DR procedures. Assess the current state, define the target, plan the work and drive it to completion together with the DevOps Lead
- Make documentation and procedure updates a non-negotiable part of delivering any change
- Replace fire-fighting with operations: planned maintenance, controlled changes, tracked problems, measured outcomes
Own incident management and troubleshooting
- Act as incident manager for major platform incidents: coordinate troubleshooting across platform, middleware, application and infrastructure teams;
manage stakeholder communication;
run post-incident reviews and drive problem management to permanent fixes - Participate in the 24/7 on-call rota
Monitoring, alerting and application management
- Ensure monitoring coverage, dashboards and alerting are meaningful: correct thresholds, clear ownership, routed to the right on-call, low noise
- Oversee application management on the platform: releases and deployments, environment management, upgrades of platform components, vulnerability and patch follow-up
Control and report
- Define and track KPIs (availability, incident volume and MTTR, change success rate, backlog age, documentation coverage, alert quality) and report regularly to management
- Maintain the platform risk register and follow mitigations through
Must-have experience and skills
- 7+ years in IT operations, application operations or platform operations, including 3+ years