Thinking Machines Lab Logo

Thinking Machines Lab

Software Engineer, Inference

Posted 5 Days Ago
In-Office
San Francisco, CA, USA
300K-350K Annually
Entry level
In-Office
San Francisco, CA, USA
300K-350K Annually
Entry level
Operate and scale production inference systems serving live, multi-tenant traffic. Own model rollouts, observability, capacity planning, incident response, reliability, failover, and cost optimization. Partner with research and inference teams to productionize serving techniques while maintaining system performance and resilience.
The summary above was generated by AI
About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up — powering Tinker's live, multi-tenant serving and the products built on top of our models.

This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online.

What You'll Do
  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform

  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production

  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly

  • Partner with inference and research teams to productionize new serving techniques without compromising reliability

  • Lead incident response for production inference issues, driving root cause analysis and durable fixes

  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows

  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale

Skills & QualificationsMinimum Qualifications
  • Experience operating large-scale, latency-sensitive production systems

  • Proficiency in Python and Go or another systems language

  • Experience with observability, monitoring, and incident response for production services

  • Strong understanding of distributed systems and how they fail at scale

Preferred Qualifications
  • Experience running production inference for large language models or other large-scale ML systems

  • Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags

  • Experience with capacity planning and cost optimization for GPU or TPU infrastructure

  • Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications

  • Comfortable being on-call and leading incident response for critical production systems

  • Comfortable working with high autonomy in a fast-changing, early-stage environment

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

Yesterday
Hybrid
San Francisco, CA, USA
180K-360K Annually
Entry level
180K-360K Annually
Entry level
Software
Develop and productionize high-performance LLM inference techniques across runtimes, scheduling, serving, routing, and GPU kernels. Profile and optimize latency, throughput, memory usage, and cost; support new models and hardware; build benchmarking frameworks; and contribute to open-source inference engines such as vLLM, SGLang, and TensorRT-LLM. The role requires strong programming, GPU architecture, ML library, and LLM optimization knowledge.
Top Skills: AwqC++CudaCutlassFp4Fp8GptqLoraPythonPyTorchSglangTensorrtTensorrt-LlmTritonVllm
Yesterday
Hybrid
San Francisco, CA, USA
180K-360K Annually
Mid level
180K-360K Annually
Mid level
Software
Build and operate the distributed inference platform powering large-scale LLM deployments. Responsibilities include orchestration, routing, autoscaling, scheduling, Model APIs, authentication, quotas, metering, observability, benchmarking, and production reliability across Kubernetes, networking, distributed runtimes, and GPU workloads. The role owns projects end to end and partners with performance engineers to deliver efficient, scalable AI inference capabilities and strong developer experiences.
Top Skills: Api GatewaysCi/Cd SystemsGpu WorkloadsKubernetesNvidia DynamoObservability ToolingService MeshesSglangTensorrt-LlmTgiVllm
2 Months Ago
Hybrid
150K-350K Annually
Senior level
150K-350K Annually
Senior level
Artificial Intelligence • Information Technology • Software
Develop and optimize inference runtimes for on-device and cloud deployments. Integrate inference engines, bring up new models and modalities, improve latency, throughput, memory, and reliability across CPU and GPU targets, build batching/scheduling/caching/distributed execution capabilities, benchmark and diagnose performance and correctness, and contribute upstream to open-source projects.
Top Skills: C++CpuCudaExecutorchGpuLlama.CppMetalMlxPythonPyTorchRocmSglangTensorrt-LlmVllmVulkan

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account