Volver a empleos
ScovaiScovaiJobs
Astra North Infoteck Inc.

Astra North Infoteck Inc.

Site Reliability Engineer – Azure Operations, Databricks Support & Incident Response

Toronto, CAPresencialContratoTiempo completo

Publicado 7 oct 2026

Este empleo está publicado en EN

Site Reliability Engineer – Azure Operations, Databricks Support & Incident Response

Waterloo or Toronto - Hybrid (Tuesday, Wednesday and Thursday)

Job Description

As an Intermediate Site Reliability Engineer, you will support and continuously improve enterprise Azure and Databricks platforms. The role focuses on production reliability, monitoring, incident response, availability, operational readiness and platform support. You will work with platform engineering, security, network, application and data teams to keep services stable, secure and supportable.

Key Responsibilities

  • Monitor and support production Azure and Databricks environments, ensuring availability, performance and operational readiness.
  • Respond to incidents and service requests and participate in on-call rotations.
  • Troubleshoot Azure platform, Databricks, networking, storage, identity, access and application-related issues.
  • Support Databricks workspaces, compute, cluster policies, jobs, workflows, user access, monitoring and cost controls.
  • Support Unity Catalog operations, including catalogs, schemas, permissions, storage credentials and external locations.
  • Support integrations between Databricks, Azure Data Lake Storage Gen2, Azure Data Factory, Azure SQL, Key Vault and managed identities.
  • Support Azure networking and connectivity components, including VNets, NSGs, routes, private endpoints, DNS, VPN/ExpressRoute and hub-and-spoke connectivity.
  • Support Azure Storage services, including storage accounts, Blob Storage and ADLS Gen2, with appropriate access, availability and lifecycle controls.
  • Manage alerts, dashboards and operational monitoring using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog or New Relic.
  • Participate in incident bridges, communicate status, follow escalation processes and support emergency changes when required.
  • Contribute to root cause analysis, problem management and follow-up remediation for recurring incidents.
  • Assist with patching, upgrades, maintenance windows, service validation and disaster recovery exercises.
  • Maintain operational runbooks, knowledge articles, support procedures and platform documentation.
  • Manage work through JIRA and ServiceNow and contribute to daily standups, planning sessions and service reviews.
  • Work with engineering teams to improve reliability, reduce recurring issues and strengthen operational support.

Candidate Requirements / Must-Have Skills

  • 3+ years of experience supporting Azure cloud infrastructure in production environments.
  • 1+ years of hands-on experience supporting Databricks environments.
  • Experience supporting Azure Storage services, including ADLS Gen2 and Blob Storage.
  • Experience working with Azure networking concepts, including VNets, NSGs, private endpoints, DNS and routing.
  • Understanding of Azure identity and security, including Entra ID, RBAC, managed identities and Key Vault.
  • Hands-on experience with monitoring and observability platforms such as Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog or New Relic.
  • Experience troubleshooting production incidents and participating in on-call support rotations.
  • Experience following incident escalation, change management, root cause analysis and problem management processes.
  • Experience working with JIRA, ServiceNow, operational runbooks and enterprise support procedures.
  • 1+ years of Windows Server administration experience.
  • 1+ years of Linux administration experience.
  • Basic understanding of network troubleshooting and connectivity concepts.

Nice-to-Have Skills

  • Experience supporting Databricks jobs, clusters, workflows, workspaces and user access.
  • Understanding of Unity Catalog concepts, permissions and governance.
  • 1+ years of Azure SQL operational support experience.
  • 1+ years of Azure Data Factory operational support experience, including linked services, integration runtimes and Databricks orchestration.
  • Understanding of Business Continuity and Disaster Recovery concepts, including RTO and RPO.
  • Experience supporting AI/GenAI platforms, Azure OpenAI, model endpoints, RAG services or MLOps operations.
  • Experience supporting enterprise data and analytics platforms.
  • Experience with capacity review, platform health reporting, cost monitoring and performance troubleshooting.
  • Strong communication, knowledge-sharing and documentation skills.

Soft Skills

  • Strong troubleshooting and analytical skills.
  • Ability to respond calmly and effectively to production incidents.
  • Ability to work independently while escalating issues appropriately.
  • Strong collaboration skills and the ability to work with engineering, platform, network, security and support teams.
  • Strong organizational skills and the ability to manage work through JIRA and ServiceNow.
  • Proactive, open-minded and willing to learn new technologies and operating models.
  • Customer-focused mindset with a commitment to reliability and operational excellence.
  • Ability to explain technical issues clearly and maintain accurate documentation.

Education

Bachelor's degree in Computer Science, Engineering, Information Technology or a related field, or equivalent practical experience.



Resumen del puesto

Tipo de empleo

Tiempo completo

Habilidades requeridas

Azure cloud infrastructure administrationDatabricks platform administration and support (workspaces, clusters, jobs, workflows)Unity Catalog operations and governance (catalogs, schemas, permissions, storage credentials)Azure Storage and ADLS Gen2 (Blob Storage, access and lifecycle controls)Azure networking and connectivity (VNets, NSGs, private endpoints, DNS, routing, VPN/ExpressRoute, hub-and-spoke)Azure identity and security (Entra ID, RBAC, managed identities, Key Vault)Monitoring and observability platforms (Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, New Relic)Incident response, on-call support and escalation processesRoot cause analysis and problem managementWindows Server administrationLinux administrationWork management and ITSM tools (JIRA, ServiceNow)Network troubleshooting and connectivity conceptsAzure Data Factory operational support (linked services, integration runtimes, Databricks orchestration)Business continuity and disaster recovery concepts (RTO, RPO)

Empleos similares

Astra North Infoteck Inc.

Open Liberty / WebSphere Liberty Lead Application Modernization Architect

Astra North Infoteck Inc.

Toronto, CAPresencialContratoTiempo completo
hace 5 horas
Astra North Infoteck Inc.

GCP Infrastructure & Cloud SRE Engineer

Astra North Infoteck Inc.

Toronto, CAPresencialContratoTiempo completo
hace 8 horas
Astra North Infoteck Inc.

AWS Sagemaker / MLOPS Engineer

Astra North Infoteck Inc.

Toronto, CAPresencialContratoTiempo completo
hace 8 horas
Astra North Infoteck Inc.

Java Full Stack Developer

Astra North Infoteck Inc.

Toronto, CAPresencialContratoTiempo completo
hace 8 horas
Maarut

RQ11491 - Product Manager - Senior

Maarut

Toronto, CAPresencialContratoTiempo completo
hace 8 horas
Maarut

RQ11648 - Business Analyst - Senior

Maarut

Toronto, CAPresencialContratoTiempo completo
hace 8 horas
Site Reliability Engineer – Azure Operations, Databricks Support & Incident Response en Astra North Infoteck Inc. en Toronto | Scovai | Scovai