SYSPRO
SRE Platform Engineer (Microsoft Azure)
Sandton, ZAPresencialPermanenteTempo integral
Publicado 17 de set. de 2026
Esta vaga foi publicada em EN
The SRE Platform Engineer is responsible for designing, building, and maintaining the reliability, availability, and scalability of Syspro’s cloud platform infrastructure. Applying site reliability engineering principles alongside modern AIOps tooling, the role ensures platform resilience, automates operational workflows, and partners with development teams to embed reliability into the software delivery lifecycle.
Beneficial Qualifications
Certified Kubernetes Administrator (CKA); AWS, GCP, or Azure cloud certifications; Google SRE or DORA-related training advantageous.
Minimum Experience
Beneficial Experience
Responsibilities
- Design and implement scalable, reliable platform infrastructure using SRE principles — including service-level objectives (SLOs), error budgets, and reliability targets — to drive consistent uptime and performance.
- Automate operational workflows through infrastructure-as-code (IaC) tooling and CI/CD pipeline configuration, eliminating manual toil and accelerating deployment velocity.
- Monitor and observe platform health across services using observability tooling (Prometheus, Grafana, or Datadog equivalents), responding to incidents and driving root-cause analysis to prevent recurrence.
- Maintain and improve Kubernetes-based container orchestration environments, ensuring optimal resource allocation, availability, and security posture across all workloads.
- Leverage AI-assisted monitoring, AIOps platforms, and ML-based anomaly detection tools to proactively identify platform risks, reduce mean time to resolution (MTTR), and surface predictive insights for the team.
- Collaborate with development teams and the Senior Kubernetes Architect to embed reliability requirements into system design from the outset, participating in design reviews and pre-production readiness assessments.
- Participate in on-call rotations and incident response, contributing to post-incident reviews and implementing corrective actions to strengthen platform resilience.
Requirements
Minimum Level of Education- Bachelor’s Degree in Computer Science, Information Technology, or a related field; or equivalent NQF Level 7 qualification.
Beneficial Qualifications
Certified Kubernetes Administrator (CKA); AWS, GCP, or Azure cloud certifications; Google SRE or DORA-related training advantageous.
Minimum Experience
- 4–6 years in a platform engineering, DevOps, or site reliability engineering role. Demonstrated hands-on experience with Kubernetes, CI/CD pipelines, and cloud infrastructure management. Proven ability to own and resolve production incidents independently.
Beneficial Experience
- Experience with multitenancy cloud platform environments. Exposure to AIOps or AI-assisted observability platforms. Terraform or Ansible infrastructure-as-code experience. Background in B2B SaaS or ERP software environments.
- Site reliability engineering: working knowledge of SLOs, SLIs, error budgets, and reliability engineering practices.
- Kubernetes and container orchestration: hands-on experience managing Kubernetes clusters, workloads, and networking configurations.
- CI/CD and IaC: proficiency with pipeline tooling (GitLab CI, GitHub Actions, Jenkins) and infrastructure-as-code (Terraform, Helm, Ansible).
- Observability and monitoring: competency with tools such as Prometheus, Grafana, Datadog, or equivalent platforms.
- AI and AIOps tooling: ability to apply AI-assisted monitoring and anomaly detection tools to improve platform health and reduce MTTR.
- Solid scripting languages (Python, Bash).
- Incident response: disciplined approach to on-call responsibilities, structured incident management, and blameless post-mortems.
- Collaboration and communication: ability to work effectively with development teams, architects, and operations leadership.
Benefits
- 25 annual leave days
- 30 days paid sick leave over 3-year cycle
- Hybrid working environment (3 office days as determined by the function/ manager)
- Pension 10%
- Medical aid
- Life insurance
- Income continuation protection
- Funeral Benefit
- Provident Fund Admin Fee
- Global Education Protection
- Bonus/ commission programmes
- Maternity: 6 months at half pay
- Paternity: 2 weeks full pay
- Employee Assistance Programme
- Free Barista Coffee Daily
- Gym Facilities
- Free Daily Lunch
- Free Fruit on a Tuesday & Thursday
- Free Refreshments & Snacks Daily
Resumo da função
Tipo de vaga
Tempo integral
lucricia.malunga@syspro.com
Competências necessárias
Site Reliability Engineering (SLOs, SLIs, error budgets, reliability engineering)Kubernetes administration and operations (cluster management, workloads)Kubernetes networking, security and resource allocationMicrosoft Azure cloud platform managementInfrastructure-as-Code (Terraform, Helm, Ansible)CI/CD pipeline tooling and configuration (GitLab CI, GitHub Actions, Jenkins)Observability and monitoring (Prometheus, Grafana, Datadog or equivalents)AIOps and AI-assisted monitoring / anomaly detectionIncident response, on-call management and blameless post‑incident reviewsScripting and automation (Python, Bash)Capacity planning, disaster recovery and business continuity designCollaboration and communication with development teams, architects and operations leadership
Vagas similares
AI Lead
SYSPRO
Sandton, ZAPresencialPermanenteTempo integral
há 6 dias
Head of Global Payroll
SYSPRO
Sandton, ZAPresencialPermanenteTempo integral
há 6 dias
TR
Operations Manager (AIA Inspections)
Two Roads Trading
Sandton, ZAPresencialPermanenteTempo integral
há 10 dias
Sales Enablement Support Specialist
SYSPRO
Sandton, ZAPresencialPermanenteTempo integral
há 20 dias
Software Engineer - Intermediate: Product Engineering
SYSPRO
Sandton, ZAPresencialPermanenteTempo integral
há 20 dias
Junior Kubernetes Architect (Microsoft Azure)
SYSPRO
Sandton, ZAPresencialPermanenteTempo integral
há 21 dias