Operations Manager – Digital Platform
Il y a 5 jours
Brussels, Brussels-Capital, Belgique
Atos
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
Operations Manager – Digital Platform & DevOps Operations
Location Brussels-50% onsite
Contract 6-12 months contract (potential extensions)
On-call Yes (24/7 escalation rota)
About the role
We are looking for an Operations Manager to bring process, control and an operations mindset to our DevOps and digital-platform operations team. The team runs the container platform, middleware and streaming services, and the operational DevOps of the mission-critical applications hosted on them. The services are operational 24/7. The team is technically strong and already has a DevOps Lead who provides technical leadership and keeps the team together day to day. That does not change. This role sits between management and the DevOps Lead and owns the operations layer: a single, agreed set of priorities; defined operational processes; documentation and procedures at a standard that survives staff turnover; monitoring and alerting that tell the truth; and reporting that gives management a reliable picture of platform health and risk. The goal is to move the team from reactive fire-fighting to planned, controlled operations — without losing the engineering strength and cohesion it already has. You decide what gets done first and how it is run; the DevOps Lead decides how it is engineered and keeps the team moving. Making that split work is the core of the job. This is a role for a process-minded operations manager with enough technical depth to be credible with senior engineers, and enough maturity to lead through an existing lead rather than around them. The environment Enterprise container platform (Red Hat OpenShift / Kubernetes) on public cloud (Azure), hosting business-critical applications Middleware stack: application servers, API gateways, identity and access management, secrets management, messaging Event streaming (Apache Kafka) and integration services CI/CD and GitOps tooling; infrastructure-as-code Observability stack: metrics, logs, traces, APM, dashboards and alerting 24/7 operations with on-call, formal incident management and service-level commitments
Key responsibilities
Own the operating model and the priorities Line-manage the DevOps Lead: set objectives, coach and support them; through them, ensure the team has clear roles, workload and on-call coverage Run a single intake and prioritisation process for all platform work — operational stability, platform improvements and requests from application teams — aligned to organisational priorities rather than individual preference Shield the team from unplanned demand, and say no, with reasons, when necessary Establish the division of responsibilities between operations management and technical leadership, and keep it clear Bring the process Define, implement and enforce operational processes for the platform: incident, problem, change and release management, on-call handover, capacity and patch management — with clear roles and decision rights Set the standard for operational documentation: runbooks, standard operating procedures, architecture and configuration records, DR procedures. Assess the current state, define the target, plan the work and drive it to completion together with the DevOps Lead Make documentation and procedure updates a non-negotiable part of delivering any change Replace fire-fighting with operations: planned maintenance, controlled changes, tracked problems, measured outcomes Own incident management and troubleshooting Act as incident manager for major platform incidents: coordinate troubleshooting across platform, middleware, application and infrastructure teams; manage stakeholder communication; run post-incident reviews and drive problem management to permanent fixes Participate in the 24/7 on-call rota Monitoring, alerting and application management Ensure monitoring coverage, dashboards and alerting are meaningful: correct thresholds, clear ownership, routed to the right on-call, low noise Oversee application management on the platform: releases and deployments, environment management, upgrades of platform components, vulnerability and patch follow-up Control and report Define and track KPIs (availability, incident volume and MTTR, change success rate, backlog age, documentation coverage, alert quality) and report regularly to management Maintain the platform risk register and follow mitigations through Must-have experience and skills 7+ years in IT operations, application operations or platform operations, including 3+ years in an operations-management role with responsibility for priorities, process and reporting in a 24/7 or mission-critical environment Experience managing through a technical lead or team leads: setting direction and standards for a team you do not run hands-on day to day Proven track record of introducing and embedding operational processes in a team that lacked them, with measurable results ITIL-based operations practice (incident, problem, change, release); ITIL 4 Foundation or higher preferred Overview-level technical knowledge — enough to lead troubleshooting and challenge engineers, not necessarily to do their job — across container platforms (OpenShift/Kubernetes), middleware (application servers, API gateways, IAM, secrets management), Kafka/event streaming, databases, CI/CD pipelines and Linux Application management
experience:
release and deployment management, environment control, upgrades, vulnerability remediation follow-up Hands-on familiarity with monitoring and observability tooling (e.g. Prometheus/Grafana, Elastic/ELK, APM tools) and with designing alerting and dashboards that on-call engineers actually use Incident management experience in production: coordinating multi-team troubleshooting, stakeholder communication, post-incident reviews and root-cause follow-through Strong writer of operational documentation and procedures, with a clear view of what "good" looks like Fluent English, written and spoken; clear, structured reporting to management Nice to have ITIL 4 Managing Professional; Kubernetes/OpenShift certification (CKA, Red Hat EX280) or equivalent exposure SRE practices: SLOs, error budgets, on-call engineering, blameless post-mortems ServiceNow (or comparable ITSM tooling) as process owner Experience in aviation, energy, telecom, finance, defence or other regulated, mission-critical sectors Exposure to security and compliance frameworks (ISO 27001, NIS2) Who you are Mission-critical mindset: stability comes first, and "it works on the cluster" is not the same as "it is operated" A process person who is technically credible: you can sit in a Kafka or OpenShift troubleshooting session, follow the reasoning and ask the question nobody wanted You lead through people, not around them: you can partner with an existing lead who is strong with the team, make them better at process, and still be the one who holds the priorities and the standard Team player and team manager at the same time: you respect strong engineers, and you still set the priorities and hold the line when the team pushes hard for its own agenda Structured, calm and decisive during incidents; transparent and evidence-based in reporting You know what good operational documentation, procedures and dashboards look like, and you can get a team there from a low base without stalling engineering Working conditions Participation in the 24/7 on-call rota as platform incident manager Occasional out-of-hours work for releases, upgrades and maintenance windows Location on-site-Brussels
- hybrid 50%
About the role
We are looking for an Operations Manager to bring process, control and an operations mindset to our DevOps and digital-platform operations team. The team runs the container platform, middleware and streaming services, and the operational DevOps of the mission-critical applications hosted on them. The services are operational 24/7. The team is technically strong and already has a DevOps Lead who provides technical leadership and keeps the team together day to day. That does not change. This role sits between management and the DevOps Lead and owns the operations layer: a single, agreed set of priorities; defined operational processes; documentation and procedures at a standard that survives staff turnover; monitoring and alerting that tell the truth; and reporting that gives management a reliable picture of platform health and risk. The goal is to move the team from reactive fire-fighting to planned, controlled operations — without losing the engineering strength and cohesion it already has. You decide what gets done first and how it is run; the DevOps Lead decides how it is engineered and keeps the team moving. Making that split work is the core of the job. This is a role for a process-minded operations manager with enough technical depth to be credible with senior engineers, and enough maturity to lead through an existing lead rather than around them. The environment Enterprise container platform (Red Hat OpenShift / Kubernetes) on public cloud (Azure), hosting business-critical applications Middleware stack: application servers, API gateways, identity and access management, secrets management, messaging Event streaming (Apache Kafka) and integration services CI/CD and GitOps tooling; infrastructure-as-code Observability stack: metrics, logs, traces, APM, dashboards and alerting 24/7 operations with on-call, formal incident management and service-level commitments
Key responsibilities
Own the operating model and the priorities Line-manage the DevOps Lead: set objectives, coach and support them; through them, ensure the team has clear roles, workload and on-call coverage Run a single intake and prioritisation process for all platform work — operational stability, platform improvements and requests from application teams — aligned to organisational priorities rather than individual preference Shield the team from unplanned demand, and say no, with reasons, when necessary Establish the division of responsibilities between operations management and technical leadership, and keep it clear Bring the process Define, implement and enforce operational processes for the platform: incident, problem, change and release management, on-call handover, capacity and patch management — with clear roles and decision rights Set the standard for operational documentation: runbooks, standard operating procedures, architecture and configuration records, DR procedures. Assess the current state, define the target, plan the work and drive it to completion together with the DevOps Lead Make documentation and procedure updates a non-negotiable part of delivering any change Replace fire-fighting with operations: planned maintenance, controlled changes, tracked problems, measured outcomes Own incident management and troubleshooting Act as incident manager for major platform incidents: coordinate troubleshooting across platform, middleware, application and infrastructure teams; manage stakeholder communication; run post-incident reviews and drive problem management to permanent fixes Participate in the 24/7 on-call rota Monitoring, alerting and application management Ensure monitoring coverage, dashboards and alerting are meaningful: correct thresholds, clear ownership, routed to the right on-call, low noise Oversee application management on the platform: releases and deployments, environment management, upgrades of platform components, vulnerability and patch follow-up Control and report Define and track KPIs (availability, incident volume and MTTR, change success rate, backlog age, documentation coverage, alert quality) and report regularly to management Maintain the platform risk register and follow mitigations through Must-have experience and skills 7+ years in IT operations, application operations or platform operations, including 3+ years in an operations-management role with responsibility for priorities, process and reporting in a 24/7 or mission-critical environment Experience managing through a technical lead or team leads: setting direction and standards for a team you do not run hands-on day to day Proven track record of introducing and embedding operational processes in a team that lacked them, with measurable results ITIL-based operations practice (incident, problem, change, release); ITIL 4 Foundation or higher preferred Overview-level technical knowledge — enough to lead troubleshooting and challenge engineers, not necessarily to do their job — across container platforms (OpenShift/Kubernetes), middleware (application servers, API gateways, IAM, secrets management), Kafka/event streaming, databases, CI/CD pipelines and Linux Application management
experience:
release and deployment management, environment control, upgrades, vulnerability remediation follow-up Hands-on familiarity with monitoring and observability tooling (e.g. Prometheus/Grafana, Elastic/ELK, APM tools) and with designing alerting and dashboards that on-call engineers actually use Incident management experience in production: coordinating multi-team troubleshooting, stakeholder communication, post-incident reviews and root-cause follow-through Strong writer of operational documentation and procedures, with a clear view of what "good" looks like Fluent English, written and spoken; clear, structured reporting to management Nice to have ITIL 4 Managing Professional; Kubernetes/OpenShift certification (CKA, Red Hat EX280) or equivalent exposure SRE practices: SLOs, error budgets, on-call engineering, blameless post-mortems ServiceNow (or comparable ITSM tooling) as process owner Experience in aviation, energy, telecom, finance, defence or other regulated, mission-critical sectors Exposure to security and compliance frameworks (ISO 27001, NIS2) Who you are Mission-critical mindset: stability comes first, and "it works on the cluster" is not the same as "it is operated" A process person who is technically credible: you can sit in a Kafka or OpenShift troubleshooting session, follow the reasoning and ask the question nobody wanted You lead through people, not around them: you can partner with an existing lead who is strong with the team, make them better at process, and still be the one who holds the priorities and the standard Team player and team manager at the same time: you respect strong engineers, and you still set the priorities and hold the line when the team pushes hard for its own agenda Structured, calm and decisive during incidents; transparent and evidence-based in reporting You know what good operational documentation, procedures and dashboards look like, and you can get a team there from a low base without stalling engineering Working conditions Participation in the 24/7 on-call rota as platform incident manager Occasional out-of-hours work for releases, upgrades and maintenance windows Location on-site-Brussels
- hybrid 50%