AgileEngine
Senior Infrastructure Engineer ID82562
León de los Aldama, MXفي الموقعدائمدوام كامل
تم النشر في 19 أغسطس 2026
تم نشر هذه الوظيفة باللغة ES
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems, with a strong focus on Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry. You will participate in on-call rotations, incident response, and root-cause analysis, automate infrastructure tasks using Python, Bash, or Go, and ensure system health across ESM and ECP platform environments. Experience with service mesh architectures is highly valued for ECP-focused roles.
WHAT YOU WILL DO
- Provide core system administration for the enterprise infrastructure.
- Focus on the maintenance, scaling, and operational stability of various on-premise and SaaS-hosted systems.
- Manage and scale containerized environments using Kubernetes.
- For ECP-focused roles: Lean heavily into executing monitoring and observability tasks to ensure system health.
- Participate in on-call rotations, incident response, and root-cause analysis (RCA).
- Automate repetitive infrastructure tasks using scripting and infrastructure-as-code (IaC).
MUST HAVES
- 4+ years of infrastructure management experience.
- Experience working with Kubernetes.
- Experience with monitoring, observability, and related SRE tasks.
- Familiarity with Snowflake and OpenTelemetry ecosystems.
- Strong background in Linux/Unix administration.
- Proficiency in scripting languages (e.g., Bash, Python, or Go).
- Upper-intermediate English level.
NICE TO HAVES
- For ECP SREs: Experience with Service Mesh architectures is highly ideal.
- Experience with Infrastructure as Code tools such as Terraform or Ansible.
- Familiarity with cloud platforms (AWS, GCP, or Azure).
PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems, with a strong focus on Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry. You will participate in on-call rotations, incident response, and root-cause analysis, automate infrastructure tasks using Python, Bash, or Go, and ensure system health across ESM and ECP platform environments. Experience with service mesh architectures is highly valued for ECP-focused roles.
WHAT YOU WILL DO
- Provide core system administration for the enterprise infrastructure.
- Focus on the maintenance, scaling, and operational stability of various on-premise and SaaS-hosted systems.
- Manage and scale containerized environments using Kubernetes.
- For ECP-focused roles: Lean heavily into executing monitoring and observability tasks to ensure system health.
- Participate in on-call rotations, incident response, and root-cause analysis (RCA).
- Automate repetitive infrastructure tasks using scripting and infrastructure-as-code (IaC).
MUST HAVES
- 4+ years of infrastructure management experience.
- Experience working with Kubernetes.
- Experience with monitoring, observability, and related SRE tasks.
- Familiarity with Snowflake and OpenTelemetry ecosystems.
- Strong background in Linux/Unix administration.
- Proficiency in scripting languages (e.g., Bash, Python, or Go).
- Upper-intermediate English level.
NICE TO HAVES
- For ECP SREs: Experience with Service Mesh architectures is highly ideal.
- Experience with Infrastructure as Code tools such as Terraform or Ansible.
- Familiarity with cloud platforms (AWS, GCP, or Azure).
PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location
ملخص الدور
نوع الوظيفة
دوام كامل
المهارات المطلوبة
KubernetesLinux/Unix AdministrationMonitoring and ObservabilityOpenTelemetrySnowflakeScripting (Bash, Python, or Go)Incident Response and Root Cause AnalysisOn-call OperationsInfrastructure Management (on-premise and SaaS)Infrastructure as Code (Terraform or Ansible)Service Mesh ArchitecturesCloud Platforms (AWS, GCP, or Azure)
وظائف مشابهة
Senior Frontend Engineer ID84686
AgileEngine
León de los Aldama, MXفي الموقعدائمدوام كامل
قبل 11 ساعة
Senior Backend Engineer ID84684
AgileEngine
León de los Aldama, MXفي الموقعدائمدوام كامل
قبل 11 ساعة
Lead Full Stack Engineer ID84682
AgileEngine
León de los Aldama, MXفي الموقعدائمدوام كامل
قبل 11 ساعة
Senior Quality Engineer ID84192
AgileEngine
León de los Aldama, MXفي الموقعدائمدوام كامل
قبل 4 أيام
Talent Community | Shopify Practice Lead ID79106
AgileEngine
León de los Aldama, MXفي الموقعدائمدوام كامل
قبل 4 أيام
Senior Infrastructure Engineer ID82566
AgileEngine
León de los Aldama, MXفي الموقعدائمدوام كامل
قبل 5 أيام