ServiceNow Logo

ServiceNow

Senior Machine Learning Engineer, Agent Eval Platform

Posted Yesterday
Be an Early Applicant
Hybrid
Santa Clara, CA, USA
Senior level
Hybrid
Santa Clara, CA, USA
Senior level
Build and calibrate an agent evaluation platform for multi-step LLM agent trajectories. Design rubrics, deterministic validators, LLM judges, confidence scoring, and human calibration workflows. Fine-tune judge models, analyze simulation-versus-production divergence, and develop process reward models for agent optimization. Establish versioned scenarios, simulated environments, and reliable measurement standards while maintaining production-grade Python code.
The summary above was generated by AI
Company Description

Who we are

Moveworks: the Agentic AI Assistant platform that empowers the entire workforce. 

Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt.

Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. 

In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work.

By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. Come join us!

ServiceNow: it all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.

Job Description

The Role

Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better?

That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.

This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

 

What you get to do in this role:

Judge design and calibration

  • A shared base judge with per-item rubrics expressed as configuration next to the dataset — so eval authors express intent, rather than forking a prompt per eval
  • Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy — was the clarifying question appropriate, was policy followed, was the path efficient
  • Scoring that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data
  • A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team — they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job
  • Fine-tuning a small judge model where an off-the-shelf one isn't good enough
  • Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction
  • Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit

Self-learning for the agent harness

This is where the pillar is headed, and a large part of why the seat exists.

  • A calibrated trajectory judge is, functionally, a reward model. Turning ours into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock
  • Using that signal to optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic
  • Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement
  • Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success

 

 

Qualifications

To be successful in this role you have:

  • 5+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used
  • Experience turning subjective human judgement into a measurement that holds up — one that other people, and ideally other models, can act on. This is the core of the job
  • Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain
  • Strong Python, and the discipline to ship production-grade code rather than notebooks
  • Ability to think and communicate clearly about complex problems — a large part of this job is convincing engineers that a number means what you say it means, and being right
  • A high degree of ownership and a bias toward shipping at startup pace
  • Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on

Experience in at least 3 of these:

  • LLM-as-judge or automated evaluation design, and calibrating it against human judgement
  • Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
  • Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly
  • Fine-tuning and evaluating small models: SFT, preference tuning, distillation
  • Reward modeling, RLHF/RLAIF, or process reward models
  • Agent trajectory analysis and step-level fault attribution
  • Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes

 

  

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. 

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license. 

HQ

ServiceNow Santa Clara, California, USA Office

2225 Lawson Lane, Santa Clara, CA, United States, 95054

ServiceNow Pleasanton, California, USA Office

4305 Hacienda Drive, Suite 200, Pleasanton, CA, United States, 94588

ServiceNow San Francisco, California, USA Office

101 Green Street, San Francisco, CA, United States, 94111

Similar Jobs at ServiceNow

13 Hours Ago
Hybrid
Santa Clara, CA, USA
126K-195K Annually
Junior
126K-195K Annually
Junior
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build evaluation infrastructure and observability for ServiceNow’s AI coding assistant. Responsibilities include scaling evaluation orchestration, designing benchmarks and scoring systems, analyzing model failures, benchmarking foundation models, and tracking agent telemetry such as token usage, cache efficiency, inference cost, and correctness.
Top Skills: Ai Coding AgentsC++Foundation ModelsJavaJavaScriptLarge Language ModelsServicenowShell
13 Hours Ago
Remote or Hybrid
166K-290K Annually
Senior level
166K-290K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead monetization strategy for CRM and Industry Workflow products, including pricing, packaging, prioritization, and product portfolio initiatives. Partner with product, go-to-market, and operations teams; prepare pricing proposals for committee approval; contribute to strategic planning; conduct pricing analytics, market research, conjoint studies, and customer value assessments; and present actionable recommendations to executives and cross-functional stakeholders.
Top Skills: Artificial Intelligence
13 Hours Ago
Hybrid
181K-317K Annually
Senior level
181K-317K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Architects and builds highly scalable distributed systems, data ingestion pipelines, and Data Lake platforms. Designs Kafka, Iceberg, Flink, and Spark solutions; develops reusable libraries and frameworks; optimizes JVM performance; troubleshoots production issues; and provides technical leadership for complex engineering initiatives. The role also requires cross-team collaboration, engineering best-practice enforcement, and evaluation of emerging technologies.
Top Skills: Apache FlinkApache IcebergSparkAvroDevOpsJavaJvmKafkaKafka ConnectMySQLOracleOrcParquetPostgres

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account