Own the reliability and quality of AI-powered features in production. Design prompting, retrieval, routing, tool-use strategies, evaluation frameworks, guardrails, fallbacks, and regression detection for non-deterministic systems. Diagnose hallucinations, instruction drift, and context sensitivity while balancing quality, cost, latency, and robustness. Collaborate with product and infrastructure teams to establish quality standards and ship dependable AI products.
ABOUT RETOOL
Nearly every company in the world runs on custom software for critical operations like tracking performance metrics, handling support workflows, building admin dashboards, and countless processes you might never have thought of. But most companies don't have the resources to properly invest in these tools, leading to a lot of old, clunky internal software, or worse, teams still stuck in manual and spreadsheet workflows.
AI has changed who gets to build software. The definition of "developer" now includes analysts, operators, and domain experts creating solutions directly—and the tools they reach for are multiplying by the week. That's both an opportunity and a challenge: as more people build with more AI tools, the risk of shipping ungoverned software into production grows just as fast.
At Retool, we're building the platform that makes all of it safe to ship. Build with any AI tool you want, then deploy into one place that connects to your real business data, enforces enterprise policies automatically, and lets teams create once and reuse everywhere with shared, trusted components. The cost of building software has collapsed. The cost of governing it hasn't—and that's the problem we solve.
Developers and domain experts have already automated over 100 million hours of work on our platform, freeing them to focus on creative problem-solving and strategic work that drives real business value. The people closest to the problem can now build the software to solve it, safely, and within enterprise guardrails.
Let's build the future together.
WHY WE’RE LOOKING FOR YOU
We’re building AI-native products where model behavior is part of the product, not just an implementation detail. As LLMs become more capable, the bottleneck is no longer access to models—it’s making non-deterministic systems reliable, evaluable, and trustworthy in production.
We’re hiring AI Engineers to own that problem end to end. This is not a role for someone who simply integrates AI APIs into features. It’s for engineers who take responsibility for how probabilistic systems behave over time, how quality is measured in the presence of variance, and how capabilities improve without regressing.
If model regressions, subtle behavior drift, or edge-case failures keep you up at night—and you enjoy that kind of ownership—we want to talk.
WHAT YOU’LL DO
As an AI Engineer, you’ll own model-driven behavior in production systems, working across product, infrastructure, and evaluation layers. Your work will directly shape what users experience—and how confidently the team can ship. You might:
- Own the behavior of AI-powered features across multiple product surfaces, including quality, safety, variance, and failure modes
- Design and evolve prompting, retrieval, routing, and tool-use strategies that embrace non-determinism while bounding its downside
- Build and maintain evaluation systems that measure model performance using statistical signals, distributions, and trends—not just pass/fail tests
- Detect, diagnose, and resolve non-deterministic failures such as hallucinations, partial correctness, instruction drift, or sensitivity to context changes
- Define and implement guardrails, fallbacks, and degradation paths that keep systems useful even when models behave unexpectedly
- Partner with product and infra teams to decide when probabilistic behavior is “good enough” to ship—and when it isn’t
- Influence model selection, model behavior, and tool design to balance quality, cost, latency, and robustness for real user workflows
You’ll work across the stack (e.g., TypeScript, Node.js, React), but your leverage won’t come from code volume alone—it will come from shaping runtime behavior with precision, measurement, and intent.
WHAT THIS ROLE IS (AND IS NOT)
This role is:
- Accountable for AI behavior, not just system correctness
- Grounded in evaluation, iteration, and regression prevention under non-determinism
- Comfortable designing systems where outputs vary, confidence is probabilistic, and correctness is contextual
- Focused on shipping dependable products on top of imperfect components
This role is not:
- Adding LLM calls to existing features and moving on
- Treating models as black boxes with undefined behavior
- Shipping AI features without owning their long-term reliability, drift, or user trust
THE SKILLSET YOU’LL BRING
- 6+ years of professional engineering experience, with ownership over complex systems in production
- Demonstrated experience owning AI/LLM behavior beyond basic integration, including mitigation of variance and failure modes
- Comfort reasoning about probabilistic systems and tradeoffs (quality vs. cost, recall vs. precision, speed vs. robustness)
- Experience designing or maintaining evaluation frameworks, golden datasets, regression detection, or human-in-the-loop feedback loops
- Strong product intuition—you care deeply about what “good” looks like even when outputs are non-deterministic
- Ability to operate independently in ambiguous problem spaces and set quality standards others rely on
- Strong opinions, weakly held—you iterate quickly and adjust based on evidence and observed runtime behavior
BONUS POINTS
- Experience with RAG, agentic systems, or tool-using models in production
- Familiarity with vector databases, embeddings, or retrieval pipelines
- Exposure to fine-tuning, model routing, or post-training techniques
- Experience building shared AI infrastructure used by multiple teams
- History of mentoring engineers on designing for non-determinism and evaluation-driven development
WHO YOU’LL WORK WITH
You’ll join a small, senior team focused on advancing AI capabilities across the product. You’ll collaborate closely with product engineers, infra engineers, designers, and PMs—often acting as the final owner of AI behavior and quality before features reach users.
Your work will set standards that others build on. If you enjoy being the person teams rely on when AI behavior matters most—and certainty is never guaranteed—you’ll thrive here.
READY TO BUILD RELIABLE AI SYSTEMS?
If you’re excited to move beyond demos and take real ownership of non-deterministic behavior in production—defining quality, preventing regressions, and turning variability into a strength—we’d love to meet you.
For candidates based in the United States, the pay range(s) for this role is listed below and represents base salary range for non-commissionable roles or on-target earnings (OTE) for commissionable roles. This salary range may be inclusive of several career levels at Retool and will be narrowed during the interview process based on a number of factors such as (but not limited to), scope and responsibilities, the candidate’s experience and qualifications, and location.
Additional compensation in the form(s) of equity and/or commission are dependent on the position offered. Retool provides a comprehensive benefit plan, including medical, dental, vision, and 401(k). Pay and benefits are subject to change at any time, consistent with the terms of any applicable compensation or benefit plans.
The base pay range for this role is $163,800 – $306,000 per year.
Retool offers generous benefits to all employees and hybrid work location. For more information, please visit the benefits and perks section of our careers page!
Retool is currently set up to employ all roles in the US and specific roles in the UK. To find roles that can be employed in the UK, please refer to our careers page and review the indicated locations.
Retool San Francisco, California, USA Office
Retool's headquarters is in San Francisco's Mission District within walking distance to the 16th St. Bart station, great coffee shops, and SF institutions like Dandelion Chocolate and Tartine Manufacturing. We have dedicated parking, secure bike storage, and 24/7 onsite security.
Similar Jobs
Financial Services
Build and operate production-grade platform services for LLM-powered agents. Responsibilities include developing reusable AI components, implementing evaluation and regression testing, creating observability capabilities, and applying security, governance, and safety controls. The role partners with product, engineering, data, security, and compliance teams to deliver scalable solutions, define success metrics, document best practices, and continuously improve reliability and business impact.
Top Skills:
Amazon Web Services (Aws)Continuous Integration/Continuous Delivery (Ci/Cd)DatabricksDockerEmbedding PipelinesKubernetesLanggraphLarge Language Models (Llms)LlamaindexPythonRetrieval-Augmented Generation (Rag)Vector Databases
Big Data • Information Technology • Software • Analytics • Energy
Contract geologist supporting regional structural and sequence-stratigraphic interpretation projects, primarily in the Anadarko Basin and other North American basins. Integrates seismic, well, production, geological, geophysical, and petrophysical data using subsurface interpretation software. Responsibilities include geological modeling, productivity analysis, multidisciplinary collaboration, quality control, stakeholder coordination, project support, and client presentations.
Top Skills:
GisKingdom
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads global HR Technology and People Analytics strategy, roadmap, investment priorities, and modernization. Oversees Oracle HCM Cloud and related HR technology solutions, aligns enterprise architecture and data strategy, and manages high-performing teams. The role drives large-scale ERP, digital transformation, technology modernization, organizational change, and best-in-class employee experiences across a complex global organization.
Top Skills:
CrunchrAzureOracle Hcm CloudSafe Agile
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



