Volver a empleos
ScovaiScovaiJobs
Astra North Infoteck Inc.

Astra North Infoteck Inc.

Observability Engineer

Halifax, CAPresencialContratoTiempo completo

Publicado 1 oct 2026

Este empleo está publicado en EN

Observability Engineer | Kubernetes, Prometheus, Grafana, Thanos, Loki

Role Description: Observability Engineer

Enterprise Kubernetes Platform | Financial Services

ABOUT THE ROLE

We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You’ll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.

WHAT YOU’LL DO

OBSERVABILITY STACK OWNERSHIP

  • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents.
  • Manage observability deployments using GitOps principles and infrastructure as code.
  • Implement long-term metrics storage solutions with cloud object storage.
  • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments.
  • Configure distributed observability architecture spanning multiple datacenters and cloud providers.

METRICS & MONITORING

  • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications.
  • Create ServiceMonitors, PodMonitors for automated metrics collection.
  • Develop rules for intelligent alerting with minimal false positives.
  • Configure multi-cluster metrics federation and aggregation.
  • Optimize metrics cardinality, storage efficiency, and query performance.
  • Implement recording rules for pre-aggregated metrics and SLI calculations.

DASHBOARDS & VISUALIZATION

  • Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics.
  • Create reusable dashboard

Resumen del puesto

Tipo de empleo

Tiempo completo

Habilidades requeridas

PrometheusGrafanaThanosLokiKubernetes observabilityGitOpsInfrastructure as CodeLong-term metrics storage with cloud object storageMulti-cluster / multi-datacenter observability architectureServiceMonitor and PodMonitor configurationAlerting strategy and rule development (intelligent alerting, reducing false positives)Metrics cardinality optimization and query performance tuningRecording rules and SLI implementationDashboard design and Grafana visualizationAI/ML-driven observability (predictive alerting, self-healing infrastructure)

Empleos similares

Recrute Action

Print Production Associate

Recrute Action

Halifax, CAPresencialContratoTiempo completo
hace 10 horas
Astra North Infoteck Inc.

AI Platform Engineer - DevOps, Python

Astra North Infoteck Inc.

Halifax, CAPresencialContratoTiempo completo
hace 7 días
Astra North Infoteck Inc.

Terraform Developer – Azure, Terraform, Kubernetes & Cloud Infrastructure

Astra North Infoteck Inc.

Halifax, CAPresencialContratoTiempo completo
hace 8 días
Métier Plus Inc

Industrial Vacuum Operator

Métier Plus Inc

Halifax, CAPresencialPermanenteTiempo completo
hace 16 días
SereneAid

Penetration Testing Specialist

SereneAid

Halifax, CAPresencialContratoTiempo completo
hace 15 días
S

IT Associate - Filipinas 1821

SOFTGIC

Manila Hilton, PHPresencialContratoTiempo completo
hace 1 hora