ML Cloud Infrastructure Engineer
Havocai · Remote · Remoto
Es un puesto remoto.
El aviso publica el sueldo: USD 150.000 a 175.000 por año.
Lo publica Havocai y está vigente desde el 14 de septiembre de 2026.
Toca Postularme y entra con tu cuenta de Google: te llevamos al aviso en Ashby y te ayudamos a armar el CV para este puesto.
PostularmeDescripción del puesto
ABOUT US:
Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk.
Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at Havoc: All-Domain Collaborative Autonomy http://havocai.com/ .
ABOUT THE ROLE
As a Machine Learning Cloud Infrastructure Engineer, you will build and operate the infrastructure HavocAI teams use to train, evaluate, deploy, and monitor machine learning models safely and reliably.
You will develop the pipelines, services, platforms, and integrations connecting data lakes, telemetry stores, simulation environments, training workloads, cloud compute, and deployed models. Your work will enable autonomy, data, and software engineers to move efficiently from raw field and simulation data to reproducible datasets, scalable training, rigorous evaluation, and production deployment.
This role is ideal for a strong software, infrastructure, or data engineer with hands-on ML experience who enjoys building production systems at the intersection of data, models, compute, and autonomy. You should be comfortable operating in a fast-paced environment, solving ambiguous infrastructure problems, and building systems that must remain scalable, observable, secure, and reliable.
WHAT YOU’LL DO
ML INFRASTRUCTURE & PIPELINES
- Build pipelines that transform raw multi-modal data—including telemetry, imagery, video, sensor, and simulation data—into curated, versioned training datasets.
- Develop reproducible training and evaluation workflows that scale across cloud compute and GPU resources.
- Build and maintain model deployment infrastructure for packaging, serving, inference, versioning, and rollback.
- Implement experiment tracking, dataset lineage, model versioning, and other capabilities required for reproducible ML development.
- Own data schema versioning and migration across pipelines, data lakes, and services as datasets and models evolve.
CLOUD PLATFORM & INFRASTRUCTURE
- Design, build, and operate scalable AWS infrastructure using Infrastructure as Code.
- Build and maintain Kubernetes/EKS workloads and containerized environments for training, batch processing, evaluation, and model serving.
- Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment.
- Improve utilization, scalability, and cost efficiency across cloud and accelerator infrastructure.
- Build infrastructure that enables engineering teams to launch workloads safely without unnecessary operational overhead.
RELIABILITY, EVALUATION & OBSERVABILITY
- Build evaluation frameworks and regression testing for model quality, dataset integrity, and pipeline correctness.
- Establish quality and reliability signals that help determine whether models are ready for production use.
- Develop monitoring, logging, tracing, and observability across training jobs, data pipelines, and deployed models.
- Diagnose and resolve performance, scaling, reliability, and infrastructure bottlenecks.
- Maintain high standards for automation, testing, documentation, and operational readiness.
CROSS-FUNCTIONAL ENGINEERING
- Partner with Autonomy, Software, Data, Simulation, and Security teams to build ML infrastructure spanning edge data capture through cloud training and model deployment.
- Contribute to CI/CD and release processes for models, datasets, and ML pipelines.
- Translate engineering requirements into scalable platform capabilities that can support multiple teams and use cases.
- Incorporate feedback from engineers and users to continuously i
Toca Postularme y entra con tu cuenta de Google: te llevamos al aviso en Ashby y te ayudamos a armar el CV para este puesto.
PostularmePreguntas frecuentes
¿Es remoto el puesto de ML Cloud Infrastructure Engineer?
Es un puesto remoto.
¿Cuánto paga?
El aviso publica USD 150.000 a 175.000 por año.
¿Dónde se publicó este aviso?
En Ashby. DameTrabajo lo encontró ahí y te lleva a postularte en el aviso original.