Senior Site Reliability Engineer (AWS/EKS and Observability) – Latin America
About Distillery
Distillery is a software development partner that helps organizations design, build, and scale modern technology solutions. Leading companies work with us to accelerate digital initiatives, solve complex technical challenges, and bring innovative products to market faster. Working as an extension of our clients' teams, we deliver lasting business value.
Distillery is committed to diversity and inclusion. We embrace a dynamic and inclusive culture where you have the opportunity to continuously learn and grow.
About the Role
Distillery is seeking a Senior Site Reliability Engineer (Latin America based) to improve the reliability of its production platform. The platform supports an open-source AI runtime and a product website, so website availability is a core product requirement. The engineer will own reliability improvements across infrastructure, observability, incident response, and production change controls, working with application engineers to diagnose failures across the browser, API, runtime, and cloud layers.
The environment includes AWS, Amazon EKS, Terraform, GPU infrastructure, and a C++ runtime, plus Python and TypeScript extensions and a JavaScript frontend. Current ops stack uses PagerDuty; a Grafana-based observability approach is under consideration.
We're looking for someone who can stabilize an existing system before introducing larger platform changes, and who works well in an early-stage environment with incomplete documentation and changing priorities.
Responsibilities
- Operate and improve production workloads on AWS and Amazon EKS.
- Maintain infrastructure as code with Terraform.
- Design actionable monitoring for infrastructure and user-facing services.
- Build useful dashboards and alerts from service-level indicators.
- Implement synthetic, browser-level, and API-level health checks.
- Participate in a shared on-call rotation across distributed time zones.
- Lead incident triage, mitigation, communication, and follow-up.
- Write clear incident reviews focused on system and process improvements.
- Create runbooks, escalation paths, and service ownership records.
- Improve deployment safeguards, rollback procedures, and configuration validation.
- Automate repetitive operational work with tested scripts and tools.
- Diagnose Linux, container, Kubernetes, network, DNS, TLS, and HTTP failures.
- Partner with developers on application monitoring and pre-production testing.
- Test recovery procedures and identify resilience gaps.
- Support secure access, secrets handling, and audit-ready operational practices.
- Communicate risks, trade-offs, and reliability recommendations clearly.
Required Experience
- 5–7 years of direct SRE work is ideal.
- A mix of SRE, DevOps, and platform engineering experience is also valuable, including 3+ years of direct SRE work.
- Related production, infrastructure, systems, or software engineering experience is valuable when it includes reliability ownership.
- 4+ years operating AWS services in production.
- 3+ years operating Kubernetes in production, including 2+ years with Amazon EKS.
- 3+ years using Terraform, including modules, state, reviews, and safe changes.
- 3+ years working with metrics, logs, traces, dashboards, and actionable alerts.
- 2+ years defining and using service-level indicators and service-level objectives.
- 3+ years responding to production incidents and leading root-cause reviews.
- 5+ years troubleshooting Linux systems and production networks.
- 3+ years using CI/CD controls, automated validation, and safe rollback methods.
- 3+ years writing reliable automation in Python, Go, Bash, or a similar language (strong in one, working knowledge of another).
- Ability to work across infrastructure and application boundaries.
- Clear written and verbal communication in a remote environment.
- A record of ownership from problem detection through verified resolution.
Preferred Experience
- 2+ years with Grafana or a similar observability platform.
- 1+ year improving on-call operations with PagerDuty or a similar platform, including alert noise and escalation paths.
- 1+ year implementing synthetic monitoring or real-user monitoring.
- 1+ year supporting production C++ services or other high-performance runtimes.
- 2+ years supporting Python applications and TypeScript or JavaScript applications.
- 1+ year supporting GPU infrastructure or AI inference workloads.
- 1+ year supporting OCR, speech, audio, or vision workloads.
- 1+ year supporting identity systems such as Zitadel or another OIDC provider.
- 1+ year supporting SOC 2 or ISO-aligned operational controls, or one complete audit cycle.
- 2+ years in an early-stage company or another fast-changing product environment.
Success Measures
- Monitoring detects user-visible failures before customers report them.
- Alerts are actionable and have a clear owner and escalation path.
- Incident response is documented, consistent, and effective.
- Repeat incidents decrease because follow-up actions are completed.
- Production changes receive appropriate automated and pre-production validation.
- Recovery procedures are documented and tested.
- Service reliability can be measured against agreed objectives.
- Application and infrastructure teams share clear operational responsibilities.
Working Model
- Remote, distributed team. Shared on-call rotation with documented handoffs.
Why You'll Like Working Here
Join a global team committed to Distillery's core values: Unyielding Commitment, Relentless Pursuit, Courageous Ambition, and Authentic Connection.
- 100% Remote Work: Enjoy the freedom to work from anywhere while collaborating with a diverse, multinational team.
- Competitive Compensation: Generous and competitive package in USD, along with a comprehensive benefits plan.
- Flexible Hours: Create a schedule that aligns with your life and priorities.
- Home Office Setup: Receive all the hardware and software needed to succeed from home.
- Innovative Workplace: Collaborate with the global Top 1% of talent in a multicultural and dynamic environment.
- Focus on Growth: Pursue professional and personal development while contributing your unique talents to a team where you can truly shine!
