Bespoke Labs Logo

Bespoke Labs

RL Environments Engineer

Reposted 16 Days Ago
Be an Early Applicant
In-Office
Mountain View, CA, USA
Entry level
In-Office
Mountain View, CA, USA
Entry level
Build and scale pipelines that generate, grade, verify, and QA reinforcement-learning environments and agentic coding tasks. Create realistic coding worlds around production codebases, automate task creation into the thousands, develop throughput-enhancing infrastructure, and manage the full task lifecycle. Analyze failures, prevent reward hacking and grader loopholes, run workloads on GCP, and use coding agents to accelerate environment development and validation.
The summary above was generated by AI
About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.

Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.

About the Role
This is a delivery role. We want an engineer who has built the machinery that turns environment ideas into hundreds or thousands of validated agentic coding tasks, and who can do it here, fast.
You will not be studying environments in the abstract. You will build the pipelines that mass-produce them, design the complex coding worlds agents train inside, and keep pushing throughput: more environments, higher quality, less manual work per task. We will measure you on the volume and quality of environments you ship, not on papers.
The thing we care about most is whether you have done this before. If you have stood up an environment-generation pipeline, scaled agentic task creation into the hundreds or thousands, and shipped it, we want to talk.

 
What You'll Do
  • Build environment-generation pipelines. Own the systems that produce RL environments programmatically, including templating, automated grading, verification, and QA, so the team ships environments at scale instead of one at a time.

  • Create complex coding worlds. Build high-fidelity environments around real codebases, with the conventions, dependencies, tooling, and technical debt that real software actually has.

  • Scale agentic task creation to thousands. Take task generation from handfuls to hundreds and thousands of validated agentic coding tasks, with automation doing the heavy lifting.

  • Build tools that raise throughput. Find the bottlenecks in environment production and remove them. Build the internal tooling and infrastructure that makes everyone on the team faster.

  • Own the full task lifecycle. Prompt, environment, grader, running frontier models against the task, failure analysis, and iteration, until each task is rigorous, fair, and hard to game.

  • Defend quality at scale. Catch reward hacking and grader loopholes, and build the verification and standards that hold the bar as volume grows.

  • Direct coding agents heavily. Use frontier coding agents to build and validate environments faster, judging their output and catching the subtle failures.

Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.

What We're Looking For

A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce.

Experience scaling that output through automation rather than through more people doing more manual work.

Strong software engineering fundamentals and fluency in several languages that holds up in production code.

Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it.

An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that.

A clear sense of what frontier coding agents can and cannot do, and where they cut corners.

Ownership. You build, debug, and ship without much supervision.

You May Be a Good Fit If You Also
  • Have worked on RL training systems, post-training, verifiers, or tool-use harnesses

  • Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure

  • Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment

  • Have contributed to a public agentic benchmark such as Terminal-Bench

  • Have open-source work that other people depend on

What We Offer
  • Location: Mountain View, CA (Onsite)

  • Base Salary: $250,000 – $300,000 USD / year

  • Additional Comp: 25% performance-based bonus + equity

Benefits & Perks

  • Health, dental, and vision coverage

  • 401(k)

  • Daily onsite lunch provided

  • Visa sponsorship and relocation support available

  • Direct impact on how the industry trains and evaluates agents

We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.

 
HQ

Bespoke Labs Mountain View, California, USA Office

800 W El Camino Real, Mountain View, California, United States, 94040

Similar Jobs

11 Days Ago
In-Office
San Francisco, CA, USA
252K-315K Annually
Senior level
252K-315K Annually
Senior level
Artificial Intelligence • Big Data • Machine Learning
Own the technical foundation for building, packaging, executing, verifying, and scaling reinforcement learning environments. Design sandboxed execution, rollout orchestration, trajectory capture, verifier frameworks, environment versioning, and authoring tools. Build high-throughput infrastructure and reliable reward signals, instrument real applications, develop task suites, and defend against reward hacking. Provide staff-level technical leadership across engineering and research teams while shipping complex production systems.
Top Skills: AWSAzureCi/CdDockerFirecrackerGCPGoGrpoGvisorInfrastructure As CodeKubernetesMcpPpoPythonRayReactRlaifRlhfRlvrRustSglangTrlTypescriptVerlVirtual MachinesVllm
24 Days Ago
In-Office
San Francisco, CA, USA
200K-200K Annually
Junior
200K-200K Annually
Junior
Artificial Intelligence • Big Data
Design datasets, evaluation rubrics, and reward signals for RLHF/RLVR; build real and synthetic data pipelines; run experiments modeling annotator behavior; develop quantitative metrics for dataset quality, diversity, and downstream impact; partner with research teams to translate training objectives into data and evaluation specs.
One Month Ago
In-Office or Remote
6 Locations
140K-200K Annually
Mid level
140K-200K Annually
Mid level
Artificial Intelligence • Information Technology • Machine Learning
The role involves designing and maintaining reinforcement learning environments, developing containerized execution environments, managing integrations, and collaborating with data operations for AI training.
Top Skills: AWSC++DockerGCPGoGymnasiumPettingzooPythonRustVm

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account