Thinking Machines Lab Logo

Thinking Machines Lab

Site Reliability Engineer, Post Training

Posted 15 Days Ago
In-Office
San Francisco, CA, USA
350K-475K Annually
Mid level
In-Office
San Francisco, CA, USA
350K-475K Annually
Mid level
Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning systems. Debug distributed failures across accelerators, networking, storage, schedulers, and training frameworks; build monitoring, alerting, recovery, checkpointing, and scheduling tools; improve cluster utilization and fault tolerance; support production model runs through on-call rotations and postmortems; and partner closely with research teams during active training runs.
The summary above was generated by AI
About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.

What You’ll Do
  • Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completion

  • Partner directly with research teams during active model runs, embedding with them to unblock training and speed up iteration

  • Debug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root cause

  • Build monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stalling

  • Improve checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of compute

  • Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads

  • Participate in an on-call rotation supporting production model runs

  • Write postmortems and turn recurring failure patterns into permanent infrastructure fixes

Skills & QualificationsMinimum Qualifications
  • 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production

  • Track record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issues

  • Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system

  • Solid grounding in Linux systems internals and networking fundamentals

  • Comfortable owning production systems, including participating in on-call rotations

Preferred Qualifications
  • Experience operating GPU or TPU training clusters at scale

  • Familiarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloads

  • Experience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)

  • Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performance

  • Experience building observability tooling purpose-built for ML training, not just general infrastructure

  • A track record of thriving in fast-changing, research-driven environments where priorities shift with the science

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000- $350,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

5 Minutes Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
140K-170K Annually
Entry level
140K-170K Annually
Entry level
Legal Tech • Software • Generative AI
Own end-to-end intake onboarding for plaintiff law firms, including voice-agent customization, call flows, firm knowledge bases, transfer rules, scoring, phone integrations, CRM/DMS handoffs, and e-signature setup. Establish onboarding standards, document repeatable processes, train internal teams, and manage intake-focused accounts. The role requires translating legal intake expertise into configurable AI workflows while collaborating with customer success, integration engineering, and law firms.
Top Skills: AIClio GrowCRMDmsE-SignatureLead DocketRingcentralSmart AdvocateVoice Ai
13 Minutes Ago
Remote or Hybrid
United States
72-80 Annually
Senior level
72-80 Annually
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Manage go-to-market execution across the UCANZ region, coordinating cross-functional campaigns, partner initiatives, timelines, dependencies, and stakeholders. Track subscriber, engagement, revenue, and campaign performance; develop dashboards, reports, business reviews, and executive recommendations. Improve operational workflows, maintain project documentation, remove blockers, and support regional leadership across Product, Marketing, Content, Analytics, Finance, Lifecycle, Operations, and external distribution partners.
Top Skills: ExcelGoogle SheetsGoogle SlidesLookerPower BIPowerPointTableau
16 Minutes Ago
Easy Apply
In-Office or Remote
Easy Apply
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Fintech • HR Tech • Financial Services
Owns end-user consumer marketing across email, direct mail, lifecycle, and other engagement channels. Responsibilities include campaign strategy and execution, onboarding and retention programs, product launches, audience segmentation, performance reporting, testing, optimization, and cross-functional collaboration with Product, Growth, Client Success, and white-label partners.
Top Skills: AI

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account