Data and ML Infrastructure Engineer
Havocai · Remote · Remoto
Es un puesto remoto.
El aviso publica el sueldo: USD 150.000 a 185.000 por año.
Lo publica Havocai y está vigente desde el 23 de septiembre de 2026.
Toca Postularme y entra con tu cuenta de Google: te llevamos al aviso en Ashby y te ayudamos a armar el CV para este puesto.
PostularmeDescripción del puesto
ABOUT US:
Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk.
Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at Havoc: All-Domain Collaborative Autonomy http://havocai.com/ .
ABOUT THE ROLE
As a Data & ML Infrastructure Engineer, you will build the data infrastructure that enables HavocAI to develop, evaluate, and continuously improve autonomous systems.
You will own the pipelines and tooling that transform large volumes of video, imagery, telemetry, sensor data, autonomy logs, and mission data into organized, searchable, and reproducible datasets. Your work will provide Autonomy and Perception engineers with the high-quality data they need to train models, evaluate system performance, reproduce failures, and improve deployed capabilities.
A major focus of this role will be HavocAI’s internal video and telemetry data lake, including ingestion, storage, indexing, metadata, curation, quality, labeling, and dataset generation.
This is a hands-on engineering role for someone who enjoys building scalable infrastructure and turning messy real-world data into reliable engineering tools and ML-ready datasets.
WHAT YOU’LL DO
DATA INFRASTRUCTURE & PIPELINES
- Build and maintain infrastructure for video, imagery, telemetry, sensor data, autonomy logs, mission data, and field-test data.
- Own data ingestion, storage, indexing, metadata, access patterns, and lifecycle management within HavocAI’s data lake.
- Develop scalable pipelines that transform raw operational data into curated datasets for ML training, evaluation, debugging, and analysis.
- Build tools for searching, filtering, tagging, and retrieving data across platforms, missions, operating conditions, and events.
- Design infrastructure capable of handling large volumes of multimodal operational data efficiently and reliably.
DATASET CURATION & ML ENABLEMENT
- Build workflows to select, clean, label, validate, and version datasets.
- Partner with Autonomy, Perception, Software, and Field Operations teams to identify high-value data for model development and system evaluation.
- Support annotation and labeling workflows for video, imagery, tracks, telemetry, and other ML inputs.
- Develop reproducible dataset-generation workflows for training, validation, regression testing, and benchmarking.
- Integrate datasets and data infrastructure with model training, experiment tracking, evaluation, and deployment workflows.
- Support multimodal dataset construction, including synchronization and alignment across sensors and data streams.
DATA QUALITY & RELIABILITY
- Develop automated checks for missing streams, corrupted files, synchronization issues, metadata gaps, labeling errors, and pipeline failures.
- Establish standards for dataset quality, lineage, versioning, and reproducibility.
- Build monitoring and observability around critical data pipelines and infrastructure.
- Troubleshoot complex data and infrastructure issues and drive them through resolution.
- Use field data, logs, and test results to help engineering teams understand system performance and identify opportunities for improvement.
DEVELOPER TOOLS & COLLABORATION
- Build self-service tools that make operational data easier for engineers to discover, access, analyze, and use.
- Partner closely with Autonomy, Perception, Software, Simulation, Field Operations, and Program teams.
- Translate engineering and ML requirements into scalable data capabilities.
- Improve workflows for replaying, visualizing, analyzing, and comparing oper
Toca Postularme y entra con tu cuenta de Google: te llevamos al aviso en Ashby y te ayudamos a armar el CV para este puesto.
PostularmePreguntas frecuentes
¿Es remoto el puesto de Data and ML Infrastructure Engineer?
Es un puesto remoto.
¿Cuánto paga?
El aviso publica USD 150.000 a 185.000 por año.
¿Dónde se publicó este aviso?
En Ashby. DameTrabajo lo encontró ahí y te lleva a postularte en el aviso original.