Braintrust Logo

Braintrust

Eval Engineer

Sorry, this job was removed at 03:08 p.m. (PST) on Friday, May 15, 2026
Be an Early Applicant
In-Office
San Francisco, CA, USA
In-Office
San Francisco, CA, USA

Similar Jobs

One Month Ago
In-Office
2 Locations
213K-263K Annually
Senior level
213K-263K Annually
Senior level
Automotive
The Senior Software Engineer will drive the architecture of core services for evaluation workflows, mentor junior engineers, and ensure APIs meet evaluation complexities.
Top Skills: C++Python
24 Minutes Ago
Remote or Hybrid
45K-85K Annually
Junior
45K-85K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handle inbound calls and warm leads, consult customers on insurance needs, recommend appropriate Property and Casualty products and coverages, and convert prospects into policyholders. The role includes paid licensing and training, customer relationship building, sales closing, brand representation, and adherence to remote work and schedule requirements.
Top Skills: Cable InternetDsl InternetFiber InternetPcWired High-Speed Internet
25 Minutes Ago
Remote or Hybrid
United States
23-45 Hourly
Internship
23-45 Hourly
Internship
Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Summer intern assisting with IT solution design and implementation, supporting cybersecurity initiatives through security protocol analysis, and collaborating with data science teams to analyze data and generate actionable insights. The role offers hands-on experience, mentorship, and professional development within Honeywell Aerospace. Candidates must remain enrolled in a relevant degree program, maintain at least a 3.0 GPA, and meet U.S. person requirements.
Top Skills: CybersecurityData ScienceInformation Technology
About the company

Braintrust is the AI observability platform. By connecting evals and observability in one workflow, Braintrust gives builders the visibility to understand how AI behaves in production and the tools to improve it.

Teams at Notion, Stripe, Zapier, Vercel, and Ramp use Braintrust to compare models, test prompts, and catch regressions — turning production data into better AI with every release.

About the role

We’re hiring an Eval Engineer to design and run creative evaluations of new AI capabilities. Your job is to turn emerging AI ideas into measurable experiments and publish the results for the developer ecosystem.

When new models, agents, or frameworks appear, everyone has opinions about what works but few people actually test them. This role exists to change that.

You’ll design experiments that compare models, prompts, and agent architectures against real tasks. You’ll build the datasets, scoring logic, and evaluation harnesses. Then you’ll publish the results so builders understand what actually works.

This role sits at the intersection of engineering, experimentation, and technical storytelling.

What you’ll ownIndustry evals
  • Design and run evaluations of new AI capabilities

  • Compare frontier models, agent systems, and tool workflows

  • Turn emerging ideas into measurable benchmarks

Eval design
  • Define datasets, tasks, and scoring logic for experiments

  • Design realistic workloads that reflect production environments

  • Create tests that expose failure modes and edge cases

Experiment implementation
  • Build evaluation harnesses using Braintrust

  • Run comparisons across models, prompts, and agent approaches

  • Analyze traces, outputs, and failure patterns

Creative test construction
  • Invent novel ways to stress test AI systems

  • Design scenarios that break agents, prompts, and model reasoning

  • Build adversarial or complex datasets that reveal weaknesses

Technical content
  • Write technical posts explaining evaluation methodology and results

  • Share datasets and scoring logic so experiments are reproducible

  • Help establish better evaluation patterns for the industry via courses

Evaluation playbooks
  • Develop reusable eval patterns for agents, RAG systems, and LLM apps

  • Create open source reference implementations developers can adopt

  • Contribute examples and guides that help teams build better evals

What great looks like
  • You’re an engineer who likes testing systems more than building features

  • You enjoy breaking things and understanding why they fail

  • You can design experiments that isolate meaningful differences between approaches

  • You understand how LLMs, agents, and RAG systems actually work

  • You write clearly for technical audiences

  • You ship experiments quickly and iterate often

  • You care about methodology and reproducibility

  • You’re curious, creative, and opinionated about how AI should be evaluated

What you’ve done
  • Built or contributed to evaluation systems for LLM or agent applications

  • Designed experiments comparing models, prompts, or AI architectures

  • Written Python code to run tests across models or APIs

  • Built datasets or scoring logic for AI quality measurement

  • Investigated model failures or unexpected behaviors

  • Published technical blog posts, research notes, or engineering write-ups

  • Built prototypes quickly to test ideas

If you want to help the industry understand how to measure AI systems and design the evaluations everyone else learns from, this is the role.

Benefits include
  • Medical, dental, and vision insurance

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

  • AI Stipend

Equal opportunity

Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

HQ

Braintrust San Francisco, California, USA Office

San Francisco, CA, United States

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account