Own the end-to-end verification system that determines whether shopping claims are true. Design and operate judge pipelines, calibration and abstention mechanisms, arbitration rulebooks, golden sets, regression corpora, and adversarial evaluations. Scale machine verdicts across millions of claims while maintaining a high accuracy floor against human-graded exams. Establish eval-driven shipping gates, control contamination, and define the seat’s success metric and decision authority with the founder. Work hands-on with models, agents, and verification infrastructure.
You own the verdict chain that decides what we publish as true for millions of shopping claims, against merchants, bots, and models that all want to fool it.
Product.ai is the verified truth layer for shopping, the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes, its first proof at scale, is the code verification service, at around $22M in revenue and roughly 60% margins. The product is codes that actually work, proven by robots running real checkouts. Founder-owned and profitable since 2009. No outside investors. No board. Fewer than twenty operators outbuilding companies 10x our size.
Why This Role Exists
Everything we sell rests on one question. Is the claim true? Machine verdicts answer it, and they decide what we publish about a product, a price, a code. The constraint on growth is whether the machine verdicts stay right as the volume climbs. You own that, end to end: the eval science behind every automated verdict, from the judges and thresholds to the exams that grade them. You work directly with the founder.
You own the arbitration rulebooks, the rules that settle a contested verdict. You also own the structural check that keeps the truth this company ships from resting on any one person's judgment, including yours. That check is a held-out, human-graded exam the machine is measured against, plus a written authority split with the founder. You are not joining a team that does this; you are the person who does it. It is also the most durable engineering seat in the building. Every model generation absorbs another layer of surface skill, but grading a verdict no human can check faster than the machine is an open frontier that grows in value for years.
The System You'll Need to Model
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
What You Will Own
You will use the craft you already own (judge design, golden sets, calibration, regression corpora, statistical rigor) and grow into the layer above: eval-gated agent orchestration, adversarial verification at commerce scale, and adjudication design for judgments no single grader can settle.
Who You Are
You interrogate every green checkmark. A passing eval is a claim, and claims get challenged, so you ask what the test could not have caught before you ask what it confirmed. You form working models of complex systems on your own. When the model is wrong, you notice and update fast. You write clearly, because clear writing is evidence of clear thought.
You treat agents as leverage you verify rather than an oracle you trust. You can do this job by hand and prove it: hand-grade a judge, hand-build a corpus, hand-verify a claim set. That mastery is exactly what lets you trust, or reject, the verdict an agent hands back. Compute is cheap here. The expensive thing is a redo cycle.
You have probably built an eval harness another team came to depend on, a regression corpus that caught a real failure before users did, an LLM-judge pipeline where you measured the judge's bias instead of trusting it, or a golden set for a system whose output is never the same twice. Adjacent roads count: scoring systems from fraud, search relevance, or content integrity, and any other domain where ground truth was scarce and the adversary was real. We care about the artifact and the reasoning behind it far more than where you did it.
Who this isn't for. This role is wrong if you optimize for leaderboard scores or trust vendor-reported evals, because the work here is deciding what a score even means instead of chasing one. It is wrong if your instinct when output looks weak is to massage the prompt rather than redesign the verification, and wrong if you would rather publish a finding than ship a gate. It is wrong if you want tightly-scoped tickets and a lane to stay in; the scope of this seat is the whole verification surface. And it is wrong if you are comfortable shipping what an agent produced without being able to say why it is right, or letting an agent grade its own work. You'll be happiest here if you have high agency and think in corpora, and if you want your verification to be the reason an entire truth layer can be trusted.
How We Evaluate
We don't run traditional AI engineering interviews. Every stage is demonstrated performance on work-relevant tasks.
Async video screen. About 15 minutes, on your own time - it replaces the recruiter screen. We want to see how you think, not how you present. Calls with company stakeholders. Short conversations with the team. Conversation with the founder. How you model the system above, where you push back, and whether you can hold the argument live. Paid work trial. One week of paid, real verification work in our environment - live loops, live corpora, real verdicts. It is paid because your time is worth paying for. We watch how you get grounded, whether you write the spec before the build, how you verify what your agents produce, and whether your self-assessment is honest.
If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.
Compensation & Ownership
Total first-year comp: $500,000 - $600,000 (base + performance-based ownership and profit-share programs). Base: $350,000 - $420,000 - top of market for principal-level machine learning engineering. Grants are performance-based, terms discussed at the offer stage.
100% premium coverage for you and your family. The token budget is effectively unlimited, steered by ROI and never capped.
This structure is built to mint partners. Based in Santa Monica, Los Angeles - in person, five days a week. Relocation support available for the right builder.
Product.ai is the verified truth layer for shopping, the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes, its first proof at scale, is the code verification service, at around $22M in revenue and roughly 60% margins. The product is codes that actually work, proven by robots running real checkouts. Founder-owned and profitable since 2009. No outside investors. No board. Fewer than twenty operators outbuilding companies 10x our size.
Why This Role Exists
Everything we sell rests on one question. Is the claim true? Machine verdicts answer it, and they decide what we publish about a product, a price, a code. The constraint on growth is whether the machine verdicts stay right as the volume climbs. You own that, end to end: the eval science behind every automated verdict, from the judges and thresholds to the exams that grade them. You work directly with the founder.
You own the arbitration rulebooks, the rules that settle a contested verdict. You also own the structural check that keeps the truth this company ships from resting on any one person's judgment, including yours. That check is a held-out, human-graded exam the machine is measured against, plus a written authority split with the founder. You are not joining a team that does this; you are the person who does it. It is also the most durable engineering seat in the building. Every model generation absorbs another layer of surface skill, but grading a verdict no human can check faster than the machine is an open frontier that grows in value for years.
The System You'll Need to Model
- Adversarial robustness, end to end. Merchants, bots, and models all have incentives to game a verdict. A merchant wants its expired code to pass as live, and a bot wants to look like a shopper. An LLM judge grades its own model family measurably too generously, a self-preference bias you measure and correct rather than assume away. Your verifier has to be structurally harder to fool than any of them.
- Calibration with decay. A verdict is more than a yes or a no. It comes with a confidence and a shelf life. A price claim goes stale in days; a materials fact stays true for years. The pipeline has to score how sure the machine is, per claim class, and admit when it is not sure. That means calibrated confidence with principled abstention, because a confidently wrong verdict is worse than no verdict at all.
- Golden sets: grading the grader. The hardest problem in the seat is meta-evaluation, measuring a judge whose judgments no human can make faster than the machine. The answer is a held-out, human-graded exam the model cannot see, cannot train on, and cannot game. Label quality, inter-rater agreement, contamination control, refresh cadence. Designing that exam is the science.
- The verification ladder. Every claim climbs one, from unverified to machine-settled to human-graded, and each rung is cheaper than the one above and harder to fool than the one below. What survives feeds Axiomatic Intelligence, our stress-tested knowledge layer, where a proven verdict gets reused instead of re-litigated.
- Evals as the ship gate. Eval-driven development is how engineering works here. Agent loops write the code and the content; your corpora and gates decide whether a loop's output ships or gets rejected. The eval is the spec.
- Cortex, the shared AI brain. The governed substrate the company runs on and our products are built on; it answers its own questions from thousands of internal documents. Your verification keeps those answers trustworthy.
- Scale, under model churn. Hundreds of thousands of merchants, millions of claims. An eval method that works on a hundred cases and falls over at a million is a demo. And the ground moves under it. Each model generation resets what judges can do and what your exams still discriminate, so the harness gets rebuilt while it runs.
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
What You Will Own
- The judge pipeline. The machine verdicts that decide what we claim is true about shopping: the models, the thresholds, and the rulebooks behind every automated verdict. Your number is the share of verdicts machines settle. You take it from a minority to the large majority and keep a high accuracy floor against the human-graded exam.
- Calibration and abstention. How sure the machine is, per claim class, each with its own rate of decay. A judge that reports 95% confidence and is right 70% of the time is miscalibrated, and that bug is yours.
- The golden sets. You grow them and keep them clean, which means label quality, contamination control, and refreshing before they saturate. You also give them teeth. When a verdict is disputed, the exam decides, not the loudest engineer in the room.
- The accuracy floor other people now pay for. Developers and AI agents buy our verified commerce data on metered keys. Anonymous access closed August 15, and the last free codes expire September 30. A miscalibrated judge is no longer an internal quality bug. It is a defect a paying customer sees. You own the floor that makes the data sellable.
- Your seat charter. Within your first quarter you co-sign a charter for this seat. It names one machine-checkable number that proves the seat is working, and a written split of what you decide freely versus what you bring to the founder to decide.
You will use the craft you already own (judge design, golden sets, calibration, regression corpora, statistical rigor) and grow into the layer above: eval-gated agent orchestration, adversarial verification at commerce scale, and adjudication design for judgments no single grader can settle.
Who You Are
You interrogate every green checkmark. A passing eval is a claim, and claims get challenged, so you ask what the test could not have caught before you ask what it confirmed. You form working models of complex systems on your own. When the model is wrong, you notice and update fast. You write clearly, because clear writing is evidence of clear thought.
You treat agents as leverage you verify rather than an oracle you trust. You can do this job by hand and prove it: hand-grade a judge, hand-build a corpus, hand-verify a claim set. That mastery is exactly what lets you trust, or reject, the verdict an agent hands back. Compute is cheap here. The expensive thing is a redo cycle.
You have probably built an eval harness another team came to depend on, a regression corpus that caught a real failure before users did, an LLM-judge pipeline where you measured the judge's bias instead of trusting it, or a golden set for a system whose output is never the same twice. Adjacent roads count: scoring systems from fraud, search relevance, or content integrity, and any other domain where ground truth was scarce and the adversary was real. We care about the artifact and the reasoning behind it far more than where you did it.
Who this isn't for. This role is wrong if you optimize for leaderboard scores or trust vendor-reported evals, because the work here is deciding what a score even means instead of chasing one. It is wrong if your instinct when output looks weak is to massage the prompt rather than redesign the verification, and wrong if you would rather publish a finding than ship a gate. It is wrong if you want tightly-scoped tickets and a lane to stay in; the scope of this seat is the whole verification surface. And it is wrong if you are comfortable shipping what an agent produced without being able to say why it is right, or letting an agent grade its own work. You'll be happiest here if you have high agency and think in corpora, and if you want your verification to be the reason an entire truth layer can be trusted.
How We Evaluate
We don't run traditional AI engineering interviews. Every stage is demonstrated performance on work-relevant tasks.
If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.
Compensation & Ownership
Total first-year comp: $500,000 - $600,000 (base + performance-based ownership and profit-share programs). Base: $350,000 - $420,000 - top of market for principal-level machine learning engineering. Grants are performance-based, terms discussed at the offer stage.
100% premium coverage for you and your family. The token budget is effectively unlimited, steered by ROI and never capped.
This structure is built to mint partners. Based in Santa Monica, Los Angeles - in person, five days a week. Relocation support available for the right builder.
Similar Jobs at Product.ai
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Own product surfaces end to end, translating founder-framed strategy into specifications, design direction, agent-built implementation, and measurable growth experiments. Lead an incentive-driven consumer verification product, protect contribution quality against gaming, and run pre-registered ship-or-kill tests. Work directly with the founder, designer, and engineers in an AI-native environment focused on verified commerce truth, agent-powered development, and high-traffic consumer surfaces.
Top Skills:
Agentic SystemsAi AgentsAi Evaluation ArchitectureAutomated Checkout SystemsExperimentation PlatformsMachine-Readable Rubrics
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Own the end-to-end pipeline for producing four to six data studies monthly from proprietary checkout, code-testing, and shopper behavior data. Responsibilities include querying BigQuery with SQL, validating statistical methods and denominators, identifying publishable findings, conducting interviews and panels, writing studies and pitch briefs, maintaining canonical statistics, and verifying AI-generated research. The role combines computational journalism, primary reporting, editorial judgment, and data-quality oversight to produce work cited by journalists and AI answer engines.
Top Skills:
Ai AgentsBigQueryGoogle Analytics 4Google Search ConsoleSQL
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Owns Product.ai’s commercial strategy for agent-economy revenue, including API and MCP monetization, keyed developer access, merchant remediation, distribution partnerships, and affiliate-network relationships. Responsible for pricing, packaging, partner development, signed pilots, and first-quarter paid revenue. The role requires independently closing technical and platform deals, building a new commercial motion from scratch, and using AI to improve sales, research, and pipeline execution.
Top Skills:
Affiliate NetworksAIAPIsCortexMcp
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

.png)