Member of Technical Staff | Observability & Reliability
Avra · São Paulo · Remoto
Es un puesto remoto: contratan desde Brasil.
El aviso no publica el sueldo.
Por el título, buscan un perfil principal o head.
Lo publica Avra y está vigente desde el 23 de septiembre de 2026.
Toca Postularme y entra con tu cuenta de Google: te llevamos al aviso en Ashby y te ayudamos a armar el CV para este puesto.
PostularmeDescripción del puesto
ABOUT THE ROLE
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.
In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.
WHAT YOU'LL DO
- Evolve our observability stack for logs, metrics, traces, and alerting.
- Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.
- Bring telemetry into customer clusters within a model where agents only make outbound connections.
- Detect drift between the desired state and what's actually running in each environment.
- Monitor the health of our deployment and runtime agents.
- Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.
- Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.
- Reduce telemetry cost: less redundant data, more useful signal.
HOW WE MEASURE SUCCESS
- 99.9% serving availability, with incidents trending down.
- MTTR, including on-premise incidents.
- Near-zero drift between desired and actual state.
- All agents active and reporting, across every dataplane.
WHAT WE'RE LOOKING FOR
- Deep experience with OpenTelemetry and observability backends.
- Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.
- Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).
- Experience operating software in environments you don't fully control.
- Production-quality code and reviews, and a willingness to operate what you build.
NICE TO HAVE
- Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).
- GCP or GKE, AWS or EKS.
- ML multi-node/multi-cluster workloads in production.
- Financial services or regulated environments.
Toca Postularme y entra con tu cuenta de Google: te llevamos al aviso en Ashby y te ayudamos a armar el CV para este puesto.
PostularmeAvisos parecidos
Preguntas frecuentes
¿Es remoto el puesto de Member of Technical Staff | Observability & Reliability?
Es un puesto remoto: contratan desde Brasil.
¿Cuánto paga?
El aviso no publica el sueldo.
¿Dónde se publicó este aviso?
En Ashby. DameTrabajo lo encontró ahí y te lleva a postularte en el aviso original.