Operations Manager-Linux

Il y a 3 jours

Brussel Hoofdstad, Brussel Hoofdstad, Belgique Atos Temps plein 90 000 € - 120 000 € Contrat

Location Brussels-50% onsite

Contract 6-12 months contract (potential extensions)

About the role

We are looking for an Operations Manager to take over leadership of the team that runs our on-premise infrastructure: the Linux server estate, virtualisation, storage and the network elements that underpin mission-critical services. These services run 24/7 and downtime has real operational consequences.

The role has two halves. First, own and tighten day-to-day operations: incident, change and configuration control, ITSM discipline in ServiceNow, and a team that knows exactly what is running, where, and why. Second, be the operational anchor for the move of this landscape away from on-premise hosting towards VMware-based cloud platforms (Azure VMware Solution or Amazon Elastic VMware Service). The migration programme is led elsewhere; this role brings the knowledge of the estate into it, sets the conditions under which the target platform is accepted into operations, and then runs it — while keeping the current estate stable throughout.

This is a hands-on management role for someone who is meticulous by nature, treats production with respect, and considers "we don't know" an unacceptable answer about a mission-critical system.

The environment

  • Linux server estate (enterprise distributions), VMware-based virtualisation, storage and backup
  • Network elements: switching, routing, firewalls, load balancers, DNS/DHCP/NTP, network segmentation
  • ITSM on ServiceNow: incident, problem, change, configuration (CMDB), knowledge
  • Target state: hybrid, then cloud-hosted VMware (AVS / EVS) with on-premise decommissioning — delivered by a migration programme that this role supports and takes into operations

Key responsibilities

Run the platform

  • Own end-to-end availability, performance and security posture of the on-premise Linux and network infrastructure against agreed SLAs/OLAs
  • Ensure patching, hardening, backup/restore and disaster-recovery procedures are defined, scheduled, executed and evidenced — not assumed
  • Manage infrastructure lifecycle: obsolescence tracking, hardware/software refresh, vendor support contracts and licences

Own the process

  • Act as process owner for infrastructure Change, Incident, Problem and Configuration Management in ServiceNow; chair or co-chair the Change Advisory Board for infrastructure changes
  • Enforce change control: every production change has a risk assessment, test evidence, a rollback plan and an approval trail
  • Keep the CMDB accurate and complete for the infrastructure estate and make it the single source of truth
  • Own major incident management for infrastructure: incident command, communication, post-incident review, and problem management through to root cause and permanent fix

Lead the team

  • Lead, prioritise and develop the infrastructure team (Linux and network engineers): workload, on-call rota, skills and performance
  • Set clear standards for documentation, runbooks and operational procedures, and make sure they are followed

Control and report

  • Define and track operational KPIs (availability, incident volume and MTTR, change success rate, patch compliance, backup success, capacity headroom) and report regularly to managementMaintain the infrastructure risk register and drive mitigations; support security, audit and compliance reviews with evidence

Support the transition and take the target state into operations

  • Be the source of truth on the current estate for the migration programme: inventory, dependencies, configurations, constraints and operational risks — and make sure migration planning is built on it
  • Review migration and target designs from an operations standpoint: supportability, monitoring, backup/DR, access, network and security controls, ITSM and CMDB integration
  • Define and enforce operational acceptance criteria for the AVS/EVS landscape: no workload is accepted into operations without runbooks, monitoring, backup, DR evidence, CMDB entries and trained on-call
  • Prepare the team and the operating model for the target platform: skills, procedures, on-call, vendor and cloud-provider support paths
  • Support cutovers from the operations side, run hyper-care, and execute on-premise decommissioning through formal change control
  • Run hybrid operations during the transition without degrading service levels

Must-have experience and skills

  • 8+ years in infrastructure operations, including 3+ years leading an infrastructure team in a 24/7, mission-critical or regulated environment (e.g. aviation, energy/utilities, telecom, finance, defence, healthcare)
  • Strong Linux operations background (RHEL/SUSE or similar): patching