Wróć do ofert
ScovaiScovaiJobs
Mps Group

Mps Group

Senior Site Reliability Engineer

Bogota, CONa miejscuUmowaPełny etat4468 USD – 5107 USD / rok

Opublikowano 6 sie 2026

To stanowisko jest opublikowane w języku ES

About MPS Group


At MPS Group, we empower businesses with cutting-edge technology solutions and expert consulting. We proudly serve a diverse range of industries, including IT, Fintech, Airlines, Energy, and more. Our expertise spans front-end technologies (React, Angular, Vue.js), back-end technologies (Java, .NET, Node.js, Python), and comprehensive Data Solutions (Database, Data Warehousing, BI, Reporting, Analytics). Our services include custom software development, QA Testing (SDET), Automation Testing, IT staffing, business development consulting, and DevOps services. We adopt both Agile and Waterfall methodologies, offering flexibility through onshore, nearshore, and offshore delivery models. From cloud solutions (Microsoft Azure, AWS, GCP) to digital advisory and project management, we craft innovative, tailored strategies designed to meet your specific business goals and drive long-term success.


About the Role


We're looking for a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives. This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design. Note: This project involves a migration from AWS to GCP, so the candidate must have strong experience with AWS and at least some exposure to GCP.


Responsibilities

  • Design and implement reliability strategies for distributed systems running across AWS and GCP.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
  • Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.
  • Automate operational processes and reduce toil through engineering solutions.
  • Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.

Requirements

  • 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong experience supporting production systems in AWS and/or GCP environments.
  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
  • Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
  • Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Bash, or comparable languages.
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
  • Experience supporting large-scale cloud migration or modernization programs (preferred).
  • Expertise in incident management and production operations for high-availability systems (preferred).
  • Experience implementing chaos engineering or resilience testing practices (preferred).
  • Knowledge of service mesh technologies such as Istio (preferred).
  • AWS and/or GCP certifications (preferred).
  • Experience working in Agile, DevOps, or DevSecOps environments (preferred).

Benefits


Details regarding benefits, compensation packages, and perks will be discussed during the hiring process.

Przegląd stanowiska

Typ stanowiska

Pełny etat

Email

malu@mps-group.us

Wymagane umiejętności

Site Reliability Engineering (SLIs, SLOs, error budgets, reliability metrics)AWS (production systems, cloud infrastructure)Google Cloud Platform (GCP) exposure/migration experienceKubernetes (EKS, GKE) operation and troubleshootingObservability and monitoring (Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk)Infrastructure as Code (Terraform)Scripting and automation (Python, Bash)Incident management, root cause analysis, and postmortemsNetworking, distributed systems, cloud security, and performance optimizationLarge-scale cloud migration and modernization programsCapacity planning and non-functional requirements designChaos engineering and resilience testing practicesService mesh technologies (Istio)

Podobne stanowiska

BF

DevOps Engineer

Baja Foundry

Bogota, CONa miejscuUmowaPełny etat
3 godziny temu
AI

Full Stack Web Developer (.NET / Python) Colombia

Advancio, Inc

Bogota, CONa miejscuUmowaPełny etat
wczoraj
AI

Business Development & Special Projects Coordinator (Mid level LATAM)

Advancio, Inc

Bogota, CONa miejscuUmowaPełny etat
5 dni temu
AI

Business Development & Special Projects Coordinator (Junior LATAM)

Advancio, Inc

Bogota, CONa miejscuUmowaPełny etat
5 dni temu
Mindtech Company

Senior Sales Executive - MT-0592

Mindtech Company

Bogota, CONa miejscuUmowaPełny etat
23 dni temu
Mps Group

Senior Data Engineer L2 or Manager (GCP + Snowflake)

Mps Group

Bogota, CONa miejscuUmowaPełny etat6216 USD – 7224 USD / rok
28 dni temu