Foxglove Logo

Foxglove

ML Platform Engineer

Sorry, this job was removed at 03:26 p.m. (PST) on Monday, Jul 13, 2026
In-Office
San Francisco, CA, USA
In-Office
San Francisco, CA, USA

Similar Jobs

38 Minutes Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
184K-348K Annually
Senior level
184K-348K Annually
Senior level
Marketing Tech • Mobile • Software
Own and evolve Braze’s ML platform for production-scale training, deployment, serving, observability, reliability, and cost efficiency. Lead complex infrastructure initiatives, including multi-region model serving, customer-specific model pipelines, CI/CD tooling, orchestration, and cloud identity. Set technical direction, manage incidents, collaborate across teams, improve engineering quality, mentor senior engineers and data scientists, and connect platform decisions to business outcomes.
Top Skills: CeleryCi/CdCloud InfrastructureFeature StoresIamInfrastructure As CodeKafkaKubernetesMl ObservabilityMlflowMongoDBNetworkingPythonRabbitMQRayRedisRuby On Rails
6 Days Ago
Remote or Hybrid
2 Locations
219K-335K Annually
Senior level
219K-335K Annually
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Lead development of scalable, reliable continuous integration infrastructure supporting autonomous-vehicle development, machine learning training, simulation, and remote builds. Specialize in Remote Build Execution and a FUSE-based file system, enabling source-code editing and developer workflows. Design and implement productivity improvements, evaluate technologies, influence technical roadmaps, establish engineering best practices, manage technical debt, and mentor engineers while balancing business and customer priorities.
Top Skills: DockerFuseGoGoogle Cloud Platform (Gcp)KubernetesNetworkingPythonRemote Build Execution (Rbe)SshUnix/Linux
16 Days Ago
Hybrid
Palo Alto, CA, USA
133K-235K Annually
Junior
133K-235K Annually
Junior
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Build and optimize large-scale machine learning infrastructure for content retrieval and recommendation. Responsibilities include developing feature generation and serving pipelines, high-performance inference systems, cloud-based training and evaluation infrastructure, and data management systems. The role partners with ML engineers to deploy models, improve reliability and efficiency, and operate highly available distributed systems using Java, Go, C++, and Python.
Top Skills: C++Caffe2FlinkGoJavaPythonPyTorchRayScikit-LearnSparkSpark MlTensorFlow
Build the data infrastructure that powers physical AI.

Physical AI is moving from research labs into production fleets across industries. As robots scale across the real world, from factories to vehicles, to defense - every workflow from product development to deployment becomes a data problem: what happened, when, on which robot, and why?

At Foxglove, we built the unified data platform for physical AI that developer and engineering teams use to answer those questions. We help teams make vast quantities of robotics data actionable, creating the data flywheel they need to develop, test, train, deploy, and operate robots with confidence.

About the role

We're looking for a ML Platform Engineer with deep infrastructure instincts to help design, deploy, and scale the systems that power Foxglove's data platform. This is a platform-first role: you'll own the infrastructure layer that makes ML possible in production, not just the models that run on top of it.

You'll be responsible for the reliability, scalability, and performance of the ML platform itself, from inference serving and pipeline orchestration to training infrastructure and evaluation frameworks. The problems are real and urgent: petabyte-scale multimodal robotics data, high-throughput retrieval and embedding pipelines, and the internal ML flywheel that lets our team ship fast. This is a hands-on infrastructure role, not research.

What you'll do
  • Design, deploy, and operate production inference infrastructure — including model serving, autoscaling, load balancing, and cost optimization across cloud environments

  • Own the platform architecture for embedding and retrieval pipelines that power semantic search over multimodal robotics data (image, video, point cloud, and timeseries)

  • Build and maintain the training and evaluation infrastructure that enables rapid iteration on model performance — including job orchestration, experiment tracking, and dataset versioning

  • Drive cloud infrastructure decisions (AWS/GCP) that directly impact latency, throughput, reliability, and cost at scale

  • Define platform abstractions and internal tooling that let product engineers ship ML-powered features without needing to manage infrastructure themselves

  • Evaluate, integrate, and operationalize third-party ML infrastructure components; establish clear build vs. buy frameworks for the team

What we're looking for
  • Deep, hands-on experience owning production ML infrastructure: inference serving, model optimization (e.g., vLLM, Triton, TorchServe), orchestration, and cloud cost management

  • Strong foundation in distributed systems and cloud infrastructure (AWS/GCP) — you think in terms of system reliability, failure modes, and operational burden, not just model accuracy

  • Experience architecting and operating retrieval systems at scale, including vector databases (e.g., Pinecone, Lance, turbopuffer, pgvector) and embedding pipelines over large, heterogeneous datasets

  • A platform engineer's mindset: you build systems that other engineers depend on, and you take that responsibility seriously

  • Proven ability to operate with high ownership — you can make hard infrastructure tradeoffs independently and move fast without breaking things

  • Strong communication skills; you can explain infrastructure tradeoffs clearly to both ML and non-ML engineers

Bonus points
  • Familiarity with fine-tuning and domain adaptation techniques for LLMs or embedding models (i.e. SFT, PEFT)

  • Familiarity with data mining or hybrid search workflows, especially as applied in robotics autonomous vehicles, or physical AI workflows

  • Prior experience building ML platforms, evaluation frameworks, or data management tooling from the ground up

Why join Foxglove
  • Work on real robotics problems. Robot data is large, messy, multimodal, time-sensitive, and tied to physical-world behavior. The problems we work on span ingestion, indexing, search, visualization, replay, connectivity, collaboration, evaluation, and operations.

  • Build tools engineers rely on. Foxglove is used by robotics teams investigating failures, validating changes, reviewing field behavior, curating datasets, and operating production fleets. The work you do helps teams understand what their robots saw, what they did, and why they behaved the way they did.

  • High-leverage product surface area. A better query path, visualization workflow, Fleet connection, UI primitive, API, onboarding flow, or customer deployment can change how an entire robotics team works.

  • Ownership and autonomy. We’re a small team, and people at Foxglove own meaningful work end-to-end. You’ll have real influence over product direction, technical architecture, customer outcomes, and how we operate as a company.

  • Strong peers and high standards. You’ll work with people who care about correctness, performance, craft, product judgment, and building software that technical users trust under pressure.

  • A mission grounded in production software. We accelerate robotics and physical AI by building the infrastructure teams use every day to connect to robots, inspect live telemetry, manage multimodal data, replay runs, investigate failures, and improve real systems.

What we offer
  • Competitive equity grant in a Series B company.

  • Medical, dental, vision, and term life insurance coverage at 100% for employees and 75% for dependents, for U.S. full-time employees.

  • 401(k) matching up to 4%, for U.S. full-time employees.

  • 4 weeks of vacation, plus holidays and winter break.

  • All-expenses-paid company offsites 1–2× per year.

  • $300 monthly budget toward commuter benefits or building your personal workspace, depending on role/location.

Learn more about how we hire and work foxglove.dev/careers

Equal opportunity

Foxglove is an equal opportunity employer. We welcome candidates from different backgrounds, experiences, and communities, and we’re committed to building an inclusive environment for everyone.

We encourage you to apply even if you don’t meet every nice-to-have listed above. The strongest candidates often bring a mix of relevant experience, curiosity, judgment, and the ability to learn quickly.

About Foxglove

Foxglove is the data platform for Physical AI. Built for robotics teams developing real-world systems, Foxglove provides a purpose-built, modular platform to collect, organize, and learn from vast quantities of multimodal data, creating the data flywheel to safely scale from development to distributed fleets. Founded in 2021, Foxglove supports hundreds of customers across automotive, aerospace, defense, logistics, agriculture, construction, and consumer robotics to deploy the next generation of intelligent machines. Learn more at foxglove.dev.

HQ

Foxglove San Francisco, California, USA Office

San Francisco, CA, United States

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account