Site Reliability Engineer
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Senior Site Reliability Engineer (Monitoring & Observability)
We are currently supporting one of our clients in the energy and infrastructure sector in their search for a Site Reliability Engineer (SRE) to strengthen their engineering team.
This organization plays a key role in the reliability and sustainability of critical infrastructure and is actively working on building a secure, scalable and highly reliable technology landscape. You will work alongside experienced engineers and cross-functional teams to ensure high availability, performance, and observability of on-premises systems.
This is an opportunity to contribute to mission-critical systems while helping to improve monitoring practices, system reliability, and operational efficiency.
Your Responsibilities
As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and observability of infrastructure and services.
Your key responsibilities will include:
-
Designing and maintaining monitoring and observability infrastructure
-
Creating dashboards, alerts, and visualization solutions for system performance and reliability
-
Implementing distributed tracing and log aggregation solutions
-
Establishing monitoring standards and defining SLI/SLO frameworks
-
Maintaining security and compliance standards within monitoring systems
-
Automating deployment processes and configuration management
-
Collaborating closely with development teams to improve application instrumentation and monitoring
-
Supporting operational reliability through participation in on-call rotations
-
Contributing to continuous improvement of infrastructure, processes, and monitoring capabilities
Required Skills & Experience
Core Technologies
-
Strong experience with Grafana
-
Experience with Prometheus and PromQL
-
Knowledge of OpenTelemetry
-
Experience with Elasticsearch for logging or monitoring solutions
Infrastructure
-
Strong Linux system administration experience
-
Solid understanding of networking concepts
-
Experience working with on-premises infrastructure environments
Programming / Automation
-
Experience with scripting or automation using Python, Bash, or Go
Experience
-
At least 3 years of experience working with monitoring and observability platforms
-
Minimum 2 years of hands-on experience with Grafana and Prometheus in production environments
-
Proven experience managing or supporting on-premises infrastructure solutions
Security & Compliance
-
Knowledge of enterprise security practices
-
Understanding of compliance requirements within infrastructure environments
Key Deliverables
-
Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through effective monitoring
-
Full observability across infrastructure and services
-
Automated monitoring, deployment, and infrastructure management
-
Monitoring systems aligned with security and compliance standards
Additional Information
-
You will work in a collaborative engineering environment with cross-functional teams.
-
The role includes participation in 24/7 on-call rotations for incident support.
-
The position offers the opportunity to contribute to large-scale, mission-critical infrastructure systems