Senior Site Reliability Engineer (AWS/EKS and Observability) – LATIN AMERICA – AWS

Latin America

*THIS CANDIDATE NEEDS TO BE IN LATIN AMERICA

*PACIFIC TIME ZONE HIGHLY PREFERRED, if other timezone there needs to be quite a bit of flexibility in working hours as this position requires realtime problem solving and support coverage.


About the Role

Rocket Ride is seeking a Senior Site Reliability Engineer to improve the reliability of its production platform. The platform supports an open-source AI runtime and a product website, so website availability is a core product requirement.

The engineer will own reliability improvements across infrastructure, observability, incident response, and production change controls, working with application engineers to diagnose failures across the browser, API, runtime, and cloud layers.

The environment includes AWS, Amazon EKS, Terraform, GPU infrastructure, and a C++ runtime, plus Python and TypeScript extensions and a JavaScript frontend. Current ops stack uses PagerDuty; a Grafana-based observability approach is under consideration.

Good fit: someone who can stabilize an existing system before introducing larger platform changes, and who works well in an early-stage environment with incomplete documentation and changing priorities.

Working Model

Remote, distributed team. Shared on-call rotation with documented handoffs.

Initial Priorities

  1. Map critical user journeys and their infrastructure dependencies.
  2. Audit existing metrics, logs, traces, dashboards, and alerts.
  3. Add outside-in checks that detect browser and API failures.
  4. Define service-level indicators and service-level objectives for critical services.
  5. Reduce alert noise and create clear alert ownership.
  6. Establish incident severity levels, escalation paths, and response procedures.
  7. Improve pre-production validation for configuration and infrastructure changes.
  8. Identify repeat failure patterns and prioritize permanent fixes.
  9. Document the production environment, operational risks, and recovery procedures.

Responsibilities

  1. Operate and improve production workloads on AWS and Amazon EKS.
  2. Maintain infrastructure as code with Terraform.
  3. Design actionable monitoring for infrastructure and user-facing services.
  4. Build useful dashboards and alerts from service-level indicators.
  5. Implement synthetic, browser-level, and API-level health checks.
  6. Participate in a shared on-call rotation across distributed time zones.
  7. Lead incident triage, mitigation, communication, and follow-up.
  8. Write clear incident reviews focused on system and process improvements.
  9. Create runbooks, escalation paths, and service ownership records.
  10. Improve deployment safeguards, rollback procedures, and configuration validation.
  11. Automate repetitive operational work with tested scripts and tools.
  12. Diagnose Linux, container, Kubernetes, network, DNS, TLS, and HTTP failures.
  13. Partner with developers on application monitoring and pre-production testing.
  14. Test recovery procedures and identify resilience gaps.
  15. Support secure access, secrets handling, and audit-ready operational practices.
  16. Communicate risks, trade-offs, and reliability recommendations clearly.

Required Experience (targets, not hard minimums — relevant experience can overlap across items)

  • 5–7 years of direct SRE work is ideal.
  • A mix of SRE, DevOps, and platform engineering experience is also valuable, including 3+ years of direct SRE work.
  • Related production, infrastructure, systems, or software engineering experience is valuable when it includes reliability ownership.
  • 4+ years operating AWS services in production.
  • 3+ years operating Kubernetes in production, including 2+ years with Amazon EKS.
  • 3+ years using Terraform, including modules, state, reviews, and safe changes.
  • 3+ years working with metrics, logs, traces, dashboards, and actionable alerts.
  • 2+ years defining and using service-level indicators and service-level objectives.
  • 3+ years responding to production incidents and leading root-cause reviews.
  • 5+ years troubleshooting Linux systems and production networks.
  • 3+ years using CI/CD controls, automated validation, and safe rollback methods.
  • 3+ years writing reliable automation in Python, Go, Bash, or a similar language (strong in one, working knowledge of another).
  • Ability to work across infrastructure and application boundaries.
  • Clear written and verbal communication in a remote environment.
  • A record of ownership from problem detection through verified resolution.

Preferred Experience (guidelines, not hard minimums — equivalent depth can substitute for elapsed time)

  • 2+ years with Grafana or a similar observability platform.
  • 1+ year improving on-call operations with PagerDuty or a similar platform, including alert noise and escalation paths.
  • 1+ year implementing synthetic monitoring or real-user monitoring.
  • 1+ year supporting production C++ services or other high-performance runtimes.
  • 2+ years supporting Python applications and TypeScript or JavaScript applications.
  • 1+ year supporting GPU infrastructure or AI inference workloads.
  • 1+ year supporting OCR, speech, audio, or vision workloads.
  • 1+ year supporting identity systems such as Zitadel or another OIDC provider.
  • 1+ year supporting SOC 2 or ISO-aligned operational controls, or one complete audit cycle.
  • 2+ years in an early-stage company or another fast-changing product environment.

Success Measures

  • Monitoring detects user-visible failures before customers report them.
  • Alerts are actionable and have a clear owner and escalation path.
  • Incident response is documented, consistent, and effective.
  • Repeat incidents decrease because follow-up actions are completed.
  • Production changes receive appropriate automated and pre-production validation.
  • Recovery procedures are documented and tested.
  • Service reliability can be measured against agreed objectives.
  • Application and infrastructure teams share clear operational responsibilities.