Build AI Logo

Build AI

ML Engineer, Inference Optimization

Posted 6 Hours Ago
In-Office
San Francisco, CA, USA
200K-320K Annually
Entry level
In-Office
San Francisco, CA, USA
200K-320K Annually
Entry level
Optimize machine-learning inference for latency, throughput, and cost. The role profiles pipelines, improves kernels, batching, quantization, compilation, serving, and hardware utilization, while building serving and evaluation infrastructure. It partners with research and product teams to make models affordable at scale and treats compute cost as a core performance metric.
The summary above was generated by AI
About Build AI

Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.

Job Summary

Inference is about 90% of compute spend. Economics are heavily driven by inference optimization. We’re hiring someone to make inference cheaper, faster, and good enough that we can scale the data engine and the product without the GPU bill eating the company.

Key Responsibilities
  • Own inference performance: latency, throughput, and cost per unit of work (tokens, frames, or jobs)

  • Cut the 90% compute line: kernels, batching, quantization, compilation, serving, and hardware utilization

  • Profile pipelines (Nsight, PyTorch Profiler, or equivalent), find the real bottleneck, and ship the fix

  • Work with research and product so models that are accurate are also affordable to run at scale

  • Build the serving and eval path so experiments don’t hide the inference bill

  • Measure cost as a first-class metric, not an afterthought once quality is “done”

You may be a good fit if you have (Must-have qualifications)
  • Strong ML / systems engineer with real inference optimization experience (serving, compilers, CUDA/kernels, quantization, or similar)

  • Comfortable in Python and in C++ or Rust for performance-critical paths

  • You think in dollars and tokens/frames per second, not only in accuracy tables

  • Familiarity with PyTorch (or JAX) and with profiling tools

  • Comfortable in a small research team shipping under cost pressure

Strong candidates may also have experience with (Nice-to-have qualifications)
  • CUDA, kernels, compilers (TVM, MLIR, TensorRT), or quantization in production

  • You have owned GPU/accelerator cost as a first-class metric

  • Serving stacks for video or large models

  • Understanding of memory hierarchy, data movement, and low-precision compute

Benefits
  • Competitive pay

  • Medical, dental, and vision packages with generous premium coverage

  • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2k per month for those living within walking distance of the office

  • Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch and dinner in our office

  • Unlimited compute budget subject to ROI justification

  • Unlimited Codex and Claude credits

  • Travel

How we're different

Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.

We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: [email protected]

Similar Jobs

One Month Ago
In-Office
Palo Alto, CA, USA
195K-262K Annually
Senior level
195K-262K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Lead end-to-end optimization of LLM/VLM inference and training infrastructure: deploy and benchmark inference stacks, diagnose performance regressions, implement quantization and compression workflows, improve latency/throughput/GPU utilization and cost per token, and collaborate with kernel, platform, research, and customer teams to productionize solutions.
Top Skills: CudaFlashinferKserveLmcacheNvidia DynamoPythonPyTorchRay ServeSglangTensorrt-LlmTritonTriton Inference ServerVllm
One Month Ago
In-Office
Palo Alto, CA, USA
250K-350K Annually
Senior level
250K-350K Annually
Senior level
Information Technology
Lead inference acceleration and GPU-parallelism optimizations (TP, SP, PP); implement high-performance CUDA/NCCL kernels; deploy videogen and LLM models to production; collaborate with researchers, conduct code reviews, mentor engineers, and optionally improve training efficiency and resource utilization.
Top Skills: Attention OptimizationCudaDeep Learning Compiler StacksGpuLarge Language ModelsNcclPipeline ParallelismQuantizationSequence ParallelismTensor ParallelismVideo Generation Models
One Month Ago
In-Office
Mountain View, CA, USA
175K-250K Annually
Mid level
175K-250K Annually
Mid level
Artificial Intelligence • Computer Vision • Hardware • Robotics
Optimize inference performance of large multimodal foundation models across cloud and on-robot targets. Diagnose bottlenecks, apply quantization/pruning/distillation, tune kernels (CUDA/Triton), build benchmarking and regression detection, and translate research models into deployment-ready implementations.
Top Skills: CudaJaxPyTorchTensorrtTorch.CompileTorchserveTritonVllmXla

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account