HavocAI Logo

HavocAI

Data and ML Infrastructure Engineer

Posted 3 Days Ago
Remote
Hiring Remotely in USA
150K-185K Annually
Mid level
Remote
Hiring Remotely in USA
150K-185K Annually
Mid level
Build and operate scalable data infrastructure and pipelines for video, imagery, telemetry, sensor, log, and mission data. Create searchable data lakes, curated and versioned ML datasets, labeling workflows, quality checks, lineage, monitoring, and reproducible dataset-generation processes. Partner with autonomy, perception, software, simulation, and field teams to support model training, evaluation, debugging, replay, and deployment workflows.
The summary above was generated by AI
About Us:

Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk.
Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at Havoc: All-Domain Collaborative Autonomy .

About the Role

As a Data & ML Infrastructure Engineer, you will build the data infrastructure that enables HavocAI to develop, evaluate, and continuously improve autonomous systems.

You will own the pipelines and tooling that transform large volumes of video, imagery, telemetry, sensor data, autonomy logs, and mission data into organized, searchable, and reproducible datasets. Your work will provide Autonomy and Perception engineers with the high-quality data they need to train models, evaluate system performance, reproduce failures, and improve deployed capabilities.

A major focus of this role will be HavocAI’s internal video and telemetry data lake, including ingestion, storage, indexing, metadata, curation, quality, labeling, and dataset generation.

This is a hands-on engineering role for someone who enjoys building scalable infrastructure and turning messy real-world data into reliable engineering tools and ML-ready datasets.

What You’ll DoData Infrastructure & Pipelines
  • Build and maintain infrastructure for video, imagery, telemetry, sensor data, autonomy logs, mission data, and field-test data.

  • Own data ingestion, storage, indexing, metadata, access patterns, and lifecycle management within HavocAI’s data lake.

  • Develop scalable pipelines that transform raw operational data into curated datasets for ML training, evaluation, debugging, and analysis.

  • Build tools for searching, filtering, tagging, and retrieving data across platforms, missions, operating conditions, and events.

  • Design infrastructure capable of handling large volumes of multimodal operational data efficiently and reliably.

Dataset Curation & ML Enablement
  • Build workflows to select, clean, label, validate, and version datasets.

  • Partner with Autonomy, Perception, Software, and Field Operations teams to identify high-value data for model development and system evaluation.

  • Support annotation and labeling workflows for video, imagery, tracks, telemetry, and other ML inputs.

  • Develop reproducible dataset-generation workflows for training, validation, regression testing, and benchmarking.

  • Integrate datasets and data infrastructure with model training, experiment tracking, evaluation, and deployment workflows.

  • Support multimodal dataset construction, including synchronization and alignment across sensors and data streams.

Data Quality & Reliability
  • Develop automated checks for missing streams, corrupted files, synchronization issues, metadata gaps, labeling errors, and pipeline failures.

  • Establish standards for dataset quality, lineage, versioning, and reproducibility.

  • Build monitoring and observability around critical data pipelines and infrastructure.

  • Troubleshoot complex data and infrastructure issues and drive them through resolution.

  • Use field data, logs, and test results to help engineering teams understand system performance and identify opportunities for improvement.

Developer Tools & Collaboration
  • Build self-service tools that make operational data easier for engineers to discover, access, analyze, and use.

  • Partner closely with Autonomy, Perception, Software, Simulation, Field Operations, and Program teams.

  • Translate engineering and ML requirements into scalable data capabilities.

  • Improve workflows for replaying, visualizing, analyzing, and comparing operational data.

  • Maintain clear documentation, data standards, and best practices for internal data use, governance, and security.

What We’re Looking For
  • Bachelor’s degree in Computer Science, Data Science, Machine Learning, Electrical Engineering, Computer Engineering, Robotics, Applied Mathematics, or a related technical field.

  • 3+ years of experience in data engineering, ML infrastructure, data platforms, backend systems, MLOps, or related engineering roles.

  • Experience designing and operating production data pipelines for large-scale structured, semi-structured, or unstructured datasets.

  • Experience working with video, imagery, time-series telemetry, sensor data, logs, or other high-volume operational data.

  • Strong programming skills in Python and SQL.

  • Experience with cloud storage, object stores, data lakes, databases, distributed processing, or modern data platforms.

  • Familiarity with dataset versioning, metadata management, data lineage, access controls, and reproducible data workflows.

  • Strong software engineering fundamentals, including testing, reliability, maintainability, and observability.

  • Strong debugging skills and comfort working across complex data pipelines and production infrastructure.

  • Ability to operate independently and take ownership in a fast-moving engineering environment.

  • U.S. citizenship and ability to obtain and maintain a U.S. Government security clearance.

Nice to Have
  • Experience with ML infrastructure, MLOps, training pipelines, experiment tracking, model evaluation, or model registries.

  • Experience managing video, perception, telemetry, or autonomous-system datasets.

  • Experience with technologies such as S3-compatible storage, PostgreSQL, Spark, Ray, Airflow, Dagster, Kubernetes, Docker, or Kafka.

  • Experience with data catalogs, dataset versioning platforms, feature stores, or labeling tools.

  • Experience building search, replay, visualization, or analysis tools for video, telemetry, logs, or sensor data.

  • Experience supporting annotation workflows for computer vision, perception, tracking, or autonomy.

  • Familiarity with sensor synchronization, timestamp alignment, calibration metadata, log replay, or multimodal dataset construction.

  • Experience with security, access controls, auditability, and data-handling requirements in government or defense environments.

  • Experience supporting defense, robotics, autonomy, aerospace, or dual-use technology programs.

  • Active or prior security clearance.

What Success Looks Like

Within your first 12 months, you will have:

  • Built reliable pipelines that move operational data from field capture into organized and searchable storage.

  • Made HavocAI’s video, telemetry, and sensor data significantly easier for engineers to discover and use.

  • Established reproducible workflows for creating high-quality datasets for model training, evaluation, and regression testing.

  • Improved data quality, lineage, metadata, and observability across critical pipelines.

  • Enabled Autonomy and Perception teams to move more quickly from field data → insight → dataset → model improvement → deployment.

Benefits:
  • 100% Employer paid Health, Dental and Vision Insurance for you and your families

  • Life Insurance (Employer Paid)

  • Ability to participate in the companies 401k program (Matching)

  • Unlimited PTO policy with an enforced 2 week minimum

  • Equity Package

  • Work / Home Office Stipend

  • Global Entry

  • 16 Week Paid Parental Leave

  • Monthly Health and Wellness Stipend


Our Values:
  • Innovation: We are driven to break new ground. Every day presents an opportunity to challenge the status quo, think boldly, and deliver advanced solutions that transform the future of defense technology.

  • Integrity: We hold ourselves to the highest ethical standards, ensuring transparency, accountability, and trust in all our actions and partnerships.

  • Mission-Driven: We are focused on achieving impactful outcomes that align with our core mission—protecting lives through innovation.

  • Forward-Leaning: We continuously seek out new opportunities and remain at the forefront of technological advancements. We embrace change and anticipate the challenges of tomorrow with confidence and creativity.

  • Ownership of All Tasks: At HavocAI, no problem is too complex or too trivial. We believe that greatness comes from tackling the hardest challenges, but also in handling the smallest, sometimes thankless, tasks with the same level of commitment and care.

  • Servant Leadership: We lead by serving others, whether it’s supporting our employees, partners, or the broader community. Empowering those around us is key to achieving long-term success and making a lasting impact.

HavocAI is an Equal Opportunity Employer and is committed to creating an inclusive and diverse workplace. We welcome applicants from all backgrounds and do not discriminate based on race, color, religion, gender, sexual orientation, age, national origin, disability, veteran status, or any other legally protected status.

Similar Jobs

One Month Ago
Remote or Hybrid
200K-275K Annually
Expert/Leader
200K-275K Annually
Expert/Leader
Artificial Intelligence • Automotive • Machine Learning • Transportation
Lead a team building large-scale ML model evaluation and orchestration frameworks for autonomous vehicles. Own data analysis workflows, dataset creation, statistical model evaluation, regression reporting, and tool development (AWS Kubernetes orchestration, ReSimulation). Provide technical leadership, mentorship, and manage project timelines and deliverables.
Top Skills: AWSKubernetesPythonResimulation
One Month Ago
Remote or Hybrid
6 Locations
Mid level
Mid level
Artificial Intelligence • Information Technology • Software
The role involves designing scalable data pipelines for 3D, video, and sensor data, optimizing infrastructure, and productionizing ML models with researchers.
Top Skills: SparkAWSAzureDaskDvcFlyteGCPKubernetesMlflowPythonPyTorchRay
25 Minutes Ago
Easy Apply
Remote or Hybrid
New Jersey, USA
Easy Apply
Expert/Leader
Expert/Leader
Artificial Intelligence • Big Data • Cloud • Security • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Owns and accelerates Nasuni’s strategic partnership with SHI and one additional national partner. Responsibilities include joint business planning, executive alignment, partner relationship development, pipeline and revenue growth, field activation, co-marketing, enablement, forecasting, Salesforce reporting, and AI-enabled workflow improvement. The role requires influencing distributed teams without formal authority, managing national partner strategy, supporting major deals, and creating repeatable programs across commercial, enterprise, and strategic segments. Travel is approximately 30–50%, with remote work preferred near Austin or New Jersey.
Top Skills: Ai-Enabled ToolsCloudCRMCybersecurityData InfrastructureSaaSSalesforceStorage

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account