Site Reliability Engineer

Il y a 1 semaine

Brussels, Belgique We ARE Recruitment Group Temps plein

Senior Site Reliability Engineer (Monitoring & Observability)

We are currently supporting one of our clients in the energy and infrastructure sector in their search for a Site Reliability Engineer (SRE) to strengthen their engineering team.

This organization plays a key role in the reliability and sustainability of critical infrastructure and is actively working on building a secure, scalable and highly reliable technology landscape. You will work alongside experienced engineers and cross-functional teams to ensure high availability, performance, and observability of on-premises systems.

This is an opportunity to contribute to mission-critical systems while helping to improve monitoring practices, system reliability, and operational efficiency.

Your Responsibilities

As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and observability of infrastructure and services.

Your key responsibilities will include:

  • Designing and maintaining monitoring and observability infrastructure

  • Creating dashboards, alerts, and visualization solutions for system performance and reliability

  • Implementing distributed tracing and log aggregation solutions

  • Establishing monitoring standards and defining SLI/SLO frameworks

  • Maintaining security and compliance standards within monitoring systems

  • Automating deployment processes and configuration management

  • Collaborating closely with development teams to improve application instrumentation and monitoring

  • Supporting operational reliability through participation in on-call rotations

  • Contributing to continuous improvement of infrastructure, processes, and monitoring capabilities

Required Skills & Experience

Core Technologies

  • Strong experience with Grafana

  • Experience with Prometheus and PromQL

  • Knowledge of OpenTelemetry

  • Experience with Elasticsearch for logging or monitoring solutions

Infrastructure

  • Strong Linux system administration experience

  • Solid understanding of networking concepts

  • Experience working with on-premises infrastructure environments

Programming / Automation

  • Experience with scripting or automation using Python, Bash, or Go

Experience

  • At least 3 years of experience working with monitoring and observability platforms

  • Minimum 2 years of hands-on experience with Grafana and Prometheus in production environments

  • Proven experience managing or supporting on-premises infrastructure solutions

Security & Compliance

  • Knowledge of enterprise security practices

  • Understanding of compliance requirements within infrastructure environments

Key Deliverables

  • Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through effective monitoring

  • Full observability across infrastructure and services

  • Automated monitoring, deployment, and infrastructure management

  • Monitoring systems aligned with security and compliance standards

Additional Information

  • You will work in a collaborative engineering environment with cross-functional teams.

  • The role includes participation in 24/7 on-call rotations for incident support.

  • The position offers the opportunity to contribute to large-scale, mission-critical infrastructure systems