ServiceNow Logo

ServiceNow

Staff Machine Learning Engineer, Agent Eval Platform

Posted 6 Days Ago
Be an Early Applicant
Hybrid
Mountain View, CA, USA
Senior level
Hybrid
Mountain View, CA, USA
Senior level
Build Moveworks’ agent evaluation and reward-modeling platform. Design configurable rubrics, deterministic validators, LLM judges, confidence scoring, and calibration against human-labeled trajectories. Analyze simulator-to-production divergence, maintain evaluation datasets, fine-tune smaller judge models, and develop process-level reward signals for optimizing prompts, tool selection, planning, retrieval, and routing. Establish versioned scenarios and reliable measurement standards for agent performance in stateful enterprise environments.
The summary above was generated by AI
Company Description

Who we are

Moveworks: the Agentic AI Assistant platform that empowers the entire workforce. 

Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt.

Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. 

In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work.

By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. Come join us!

ServiceNow: it all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.

Job Description

The Role

Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better?

That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.

This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

 

What you get to do in this role:

Judge design and calibration

  • A shared base judge with per-item rubrics expressed as configuration next to the dataset — so eval authors express intent, rather than forking a prompt per eval
  • Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy — was the clarifying question appropriate, was policy followed, was the path efficient
  • Scoring that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data
  • A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team — they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job
  • Fine-tuning a small judge model where an off-the-shelf one isn't good enough
  • Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction
  • Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit

Self-learning for the agent harness

This is where the pillar is headed, and a large part of why the seat exists.

  • A calibrated trajectory judge is, functionally, a reward model. Turning ours into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock
  • Using that signal to optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic
  • Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement
  • Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success

 

 

Qualifications

To be successful in this role you have:

  • 8+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used
  • Experience turning subjective human judgement into a measurement that holds up — one that other people, and ideally other models, can act on. This is the core of the job
  • Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain
  • Strong Python, and the discipline to ship production-grade code rather than notebooks
  • Ability to think and communicate clearly about complex problems — a large part of this job is convincing engineers that a number means what you say it means, and being right
  • A high degree of ownership and a bias toward shipping at startup pace
  • Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on

Experience in at least 3 of these:

  • LLM-as-judge or automated evaluation design, and calibrating it against human judgement
  • Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
  • Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly
  • Fine-tuning and evaluating small models: SFT, preference tuning, distillation
  • Reward modeling, RLHF/RLAIF, or process reward models
  • Agent trajectory analysis and step-level fault attribution
  • Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes

 

  

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. 

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license. 

HQ

ServiceNow Santa Clara, California, USA Office

2225 Lawson Lane, Santa Clara, CA, United States, 95054

ServiceNow Pleasanton, California, USA Office

4305 Hacienda Drive, Suite 200, Pleasanton, CA, United States, 94588

ServiceNow San Francisco, California, USA Office

101 Green Street, San Francisco, CA, United States, 94111

Similar Jobs at ServiceNow

Yesterday
Hybrid
Santa Clara, CA, USA
241K-433K Annually
Senior level
241K-433K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads enterprise customer experience strategy at ServiceNow, owning NPS methodology, customer journey and lifecycle modeling, CX investment prioritization, and executive and board-level reporting. Influences Product, Sales, Support, Marketing, and Customer Excellence leaders to address customer friction and improve renewal, expansion, and support outcomes. Manages a three-person senior team and establishes measurement standards, accountability, and follow-through.
Yesterday
Hybrid
Santa Clara, CA, USA
176K-308K Annually
Senior level
176K-308K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design and lead production agentic AI systems for CPQ workflows, including multi-agent orchestration, async execution, streaming, state management, tool boundaries, MCP and A2A interoperability, RAG, and context engineering. Own the agent execution Harness, make architecture and performance decisions, support classical ML pipelines, and provide full-stack integration across Python, Java/Spring Boot, React, TypeScript, and Lit. Work directly with customers and product teams to deploy solutions and set technical direction.
Top Skills: A2A ProtocolAsyncioEmbeddingsFastapiGraph-Based RetrievalJavaLangchainLanggraphLitLlmsMcpPythonPyTorchRagReactScikit-LearnSpring BootTypescriptVector RetrievalWebsockets
Yesterday
Hybrid
Santa Clara, CA, USA
176K-308K Annually
Expert/Leader
176K-308K Annually
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead architecture, roadmap, and implementation for Veza’s scalable Access Graph data platform. Solve complex distributed-systems and graph-database challenges, establish engineering standards, improve reliability and performance, and deliver backend services in Go. Drive cross-team planning, technical decisions, design reviews, and stakeholder alignment while mentoring engineers and team leads. Ensure high standards for correctness, testing, documentation, support, rollback safety, and responsible use of AI-assisted engineering tools.
Top Skills: AWSAzureClaude CodeDockerGCPGithub CopilotGoGraph DatabasesKubernetesMicroservicesNeo4JOlapOltpRestful Apis

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account