Archetype AI Logo

Archetype AI

Machine Learning Systems Engineer

Posted 5 Hours Ago
Be an Early Applicant
In-Office
San Mateo, CA, USA
Senior level
In-Office
San Mateo, CA, USA
Senior level
Own the production serving path for multimodal AI models, including Rust inference runtime nodes, GPU kernels, routing, batching, streaming, quantization, observability, and low-latency optimization. Productionize research checkpoints and improve GPU utilization, numerical precision, latency, and cost while collaborating closely with researchers.
The summary above was generated by AI
About Job

At Archetype AI, we’re building the world’s first physical AI platform to bring artificial intelligence into the real world. Our foundation model, Newton, understands the physical world through objective sensor data and generates real-time insights into complex physical behaviors, from industrial machinery and systems to wearable devices and smart environments.

Formed by a high-caliber team from Google and backed by one of Silicon Valley’s most renowned venture funds, Archetype AI is in a Series A phase and rapidly advancing its technology for the next big leap. This is a unique opportunity to join an exciting, fast-growing AI team based in the heart of Silicon Valley.

About the Role

You will own the serving path for Newton and related multimodal models. Much of our inference stack is Rust-native: model nodes in our agent runtime, built on Rust ML stacks (candle, Burn) with custom GPU kernels, plus the routing layer that streams real-time inference to GPU nodes. You will drive GPU utilization, numerical precision, and low-latency serving from the kernel up.

What You'll Own
  • Build and own model nodes in our Rust inference runtime: loading, warmup, batching, streaming, GPU memory pools.

  • Optimize kernels and the GPU path: custom CUDA kernels, mixed precision, quantization, parity against research.

  • Own inference routing and serving: streaming path API to GPU node, request batching, SLO-backed latency and cost.

  • Productionize research checkpoints: export, compilation, quantization, parity evals, and rollout.

  • Build the observability inference needs: latency histograms, GPU metrics, OOM signatures, replayable traces.

Key Qualifications
  • 6+ years software engineering, several of them in ML systems, inference, or high-performance GPU computing.

  • Has owned a production serving path end to end, not only benchmarked models.

  • Expert in PyTorch with models shipped to production; knows what will be slow before the profiler runs.

  • Strong CUDA or equivalent GPU depth: memory hierarchy, occupancy, Nsight or equivalent profiling.

  • Rust or C++ alongside Python, Linux performance, production ops; ready to work in Rust daily.

  • Works with researchers: can translate an architecture change into serving work.

Nice to Have
  • Rust ML stacks: candle, Burn, or comparable GPU compute in Rust.

  • Custom kernels and compiler stacks: Triton, CUTLASS, TorchInductor, TensorRT.

  • Quantization and mixed precision in production with a numerical-correctness suite.

  • Multimodal, video, embedding or time-series serving, not only decoder-only chat LLMs.

  • High-performance serving stacks (vLLM, SGLang, TensorRT-LLM): continuous batching, paged KV cache.

Similar Jobs

2 Days Ago
In-Office or Remote
San Francisco, CA, USA
171K-269K Annually
Senior level
171K-269K Annually
Senior level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Design and optimize large-scale machine learning model-serving systems, including distributed infrastructure, load balancing, auto-scaling, batching, caching, and inference optimization. Build highly reliable, high-concurrency services serving billions of requests, improve latency and throughput, benchmark inference engines, and develop CI/CD infrastructure for model deployments. Partner with ML engineers to fine-tune and deploy open-source large language models.
Top Skills: Auto-ScalingBatch SchedulingCC++Ci/CdDistributed SystemsGpu KernelsKv CacheLoad BalancingQuantizationRustSglangSpeculative DecodingTensorrt-LlmTritonVllm
14 Days Ago
Hybrid
Mountain View, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build and operate reliable production systems for machine learning models, agentic workflows, and self-learning capabilities. Responsibilities include ML lifecycle automation, continuous delivery, safe deployments, feedback loops, observability, SLOs, distributed workload optimization, incident response, platform automation, and cloud infrastructure standards. The role provides staff-level technical leadership across ML, data, product, and platform engineering teams.
Top Skills: C++Ci/CdCloud InfrastructureContainersDistributed SystemsGoGpu ComputingInfrastructure As CodeJavaKubernetesMachine Learning OperationsObservabilityPythonRustSlis/Slos
20 Days Ago
In-Office or Remote
7 Locations
277K-415K Annually
Expert/Leader
277K-415K Annually
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Build and operate production machine learning systems for ranking, retrieval, recommendations, search, propensity, churn, LTV, and next-best-action decisioning. Design reliable signal contracts covering freshness, provenance, confidence, eligibility, and calibration. Lead feature pipelines, model serving, experimentation, monitoring, and feedback loops while evaluating fairness, risk, compliance, trust, and long-term customer impact. Collaborate across product, growth, data, platform, modeling, risk, and compliance teams.
Top Skills: Ai AgentsBatch PipelinesData LakehousesData WarehousesEmbeddingsEvent StreamsExperimentation SystemsFeature StoresJavaKotlinKubernetesLarge Language ModelsLightgbmModel-Serving InfrastructureObservability ToolingPythonPyTorchRecommendation SystemsSemantic SearchSQLTensorFlowWorkflow OrchestrationXgboost

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account