Monitoring Operator
Il y a 4 jours
Brussels, Belgique
ProUnity
Temps partiel
500 € - 580 € Contrat
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
Tasks:
- Master the entire Zabbix environment: creating and maintaining dashboards, network maps, triggers, items, and templates. Configure and adjust alert thresholds (triggers) according to business and technical needs, including via user macros and dependent triggers to limit noise.
- Create and maintain reusable monitoring templates (items, macros, LLD, low-level discovery)
- Design and maintain Zabbix Agent, SNMP, JMX, HTTP Agent and External Checks items according to use cases
- Develop custom discovery scripts (LLD) for the self-discovery of equipment and services
- Diagnosing false positives, noisy alerts, and collection anomalies (proxy, pollers, timeouts) in collaboration with the technical teams under the supervision of the manager
- Ensure daily monitoring of active alerts and their classification (actual incident vs. noise)
- Configure and maintain Zabbix ↔ ServiceNow integrations (automatic incident escalation) and Zabbix ↔ notification tools
- Administer the distributed Zabbix architecture (proxies, servers, high availability) and ensure its performance
- Participate in the evolution of the monitoring architecture (proxies, SNMP v2/v3 integrations, VMware API, NetFlow, etc.)
- Perform root cause analysis (RCA) on complex multi-layered incidents (network, system, application, storage) at the request of technical teams
- Mastering the Zabbix Services module (Business Service Monitoring): building hierarchical service trees, aggregating status (SLA/SLO) from triggers, calculating and monitoring availability rates per business service
- Modeling end-to-end application services by associating technical components (hosts, triggers) with the relevant business services provides a user/business-oriented view rather than a purely infrastructure-oriented one.
- Configure problem tags and state propagation rules for accurate mapping between technical incidents and their impact on services.
- Leverage the SLA reports generated by the Services module to feed into monthly reporting and availability reviews
Oh Dear - Website Monitoring:
- Set up and maintain technical monitoring of websites on Oh Dear (availability, SSL certificates, performance, uptime, broken links, mixed content)
- Configure alerts (email, webhooks, third-party integrations) and check the status of monitored sites daily.
- Leverage Oh Dear’s REST API for data extraction and integration with reporting tools (Power BI is a plus)
- Monitor detected incidents and escalate them to the relevant teams, including impact analysis.
ServiceNow – Ticket and Incident Management:
- Daily review of incidents created in ServiceNow related to monitoring and ensure their technical qualification
- Handle all requests from IT teams arriving via our request catalog in ServiceNow and by email
- Qualify, prioritize (based on impact/urgency) and process monitoring-related requests received via ServiceNow or email
- Contribute to the reliability of the CMDB in relation to discovery processes (VMware, network) and the resolution of identification conflicts
- Ensure rigorous follow-up until tickets are closed, with documentation of actions taken.
Main Activities:
Ensure the administration, operation, and development of monitoring and observability tools, under the supervision of the Monitoring & CMDB manager. Activities include:
- Administering Zabbix: creating and maintaining dashboards, network maps, items, triggers and templates, automating discovery and optimizing the monitoring architecture.
- Model and monitor end-to-end business and application services, including via the Zabbix Services module, to assess the impact of incidents and track availability and SLA/SLO commitments.
- Monitor websites with Oh Dear: availability, SSL certificates, performance and anomalies.
- Qualify and process alerts, incidents and monitoring requests in ServiceNow, ensure their follow-up until closure and contribute to the reliability of the CMDB.
- Improve observability by correlating metrics, logs and events, reducing unnecessary alerts and contributing to incident diagnosis with technical teams.
- Develop integrations between tools and automate controls, data collection and reporting using REST APIs and Python/Bash scripts.
- Produce availability and performance reports, analyze recurring incidents and maintain management dashboards.
- Apply Git, CI/CD and Infrastructure as Code practices, maintain technical documentation and ensure knowledge transfer to support and operations teams.