NVIDIA Logo

NVIDIA

Senior Staff Software Engineer - Enterprise AI Platform

Posted 6 Days Ago
Be an Early Applicant
In-Office
Santa Clara, CA, USA
200K-322K Annually
Senior level
In-Office
Santa Clara, CA, USA
200K-322K Annually
Senior level
Design and build NVIDIA’s enterprise AI agent platform, including agent blueprints, runtime safety and policy enforcement, multi-agent orchestration, credential brokering, checkpointing, observability, evaluation, and self-improvement loops. Deploy reliable long-running services at scale, integrate enterprise systems, support secure execution, and participate in on-call operations. The role requires extensive distributed systems or infrastructure experience and hands-on agent development.
The summary above was generated by AI

We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own.

Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill, safety policy, or trace works the same no matter which harness is running. We want to enable agents that act on a person's behalf, governed and secure, continuously evaluated and self-improving. These agents coordinate and hand work off to each other, with identity and policy following every hop. They route and tune themselves across harnesses from live eval signals, and get better from their own production telemetry instead of waiting on a human to retrain them. Have you run agents on a harness like Claude or Codex and hit the walls that show up when they run for real, for days, against live systems — and wanted them to learn from it on their own? We're building the platform that solves those problems once, for every team.

What you'll be doing:

The day-to-day is designing and building the platform that agent builders across NVIDIA depend on, from engineering teams to functions like finance, legal, and HR:

  • Design agent blueprints with clean interfaces for authorization, sandbox, memory, observability, and skills, so a new agent inherits its enterprise integrations from the platform.

  • Build the runtime safety harness: a policy engine that checks every action before it runs, rate and budget caps, circuit breakers, approval gates, action allow-lists, and a kill switch that works even when an agent goes rogue.

  • Enable composing and orchestrating agents: skills as first-class units with declarative manifests, multi-agent orchestration for delegation and handoff, and support for headless, long-running, autonomous agents.

  • Broker credentials so multi-agent systems can authenticate and authorize without ever touching secrets, with least-privilege scoping on every token.

  • Provide checkpoint and recovery so an agent resumes cleanly after a crash or restart.

  • Instrument observability and evaluation that span harnesses: decision-level traces with correlation IDs, what the agent saw and chose and why, cost anomaly alerts that catch looping, and quality scoring across skills, whole agents and products. Then close the loop: turn those signals into insights that make the agents better.

Since teams across NVIDIA depend on this platform in production, the whole team shares in keeping it healthy and reliable, including taking part in an on-call rotation.

What we need to see:

  • BS or MS in Computer Science, Engineering, or related field (or equivalent experience)

  • 12+ years building distributed systems, infrastructure, or developer platforms at scale

  • Hands-on experience building agents on a harness, exposing them as APIs, and shipping them with CI/CD

  • Experience deploying and operating long-running services on container orchestration platforms

  • Experience with the building blocks of scalable systems: messaging, caching, and durable storage

  • Proficiency in Python, Go, Rust, or similar

Ways to stand out from the crowd:

  • Built a safety or policy engine that enforces rules on agent actions at runtime, with approval gates and kill switches

  • Designed evaluation and feedback loops for agent behavior, tied to versioned skills or blueprints. Built self-evolving loops where agents improve from their own eval and production signals, on the latest agent harnesses — the closed-loop, self-improving side you want to attract

  • Applied security fundamentals like threat modeling, authentication and authorization, least privilege, secrets management, and token exchange

  • Designed AI data platform components like ingestion pipelines, vector stores, and retrieval APIs

  • Shipped platform building blocks adopted by multiple engineering teams. Led complex technical projects like migrations or greenfield platform builds, aligning teams and writing clear design docs

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 13, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Santa Clara, California, USA Office

2701 San Tomas Expressway, Santa Clara, CA, United States, Santa Clara

NVIDIA San Francisco, California, USA Office

San Francisco, United States

NVIDIA San Jose, California, USA Office

San Jose, United States

Similar Jobs

A Minute Ago
Hybrid
San Francisco, CA, USA
244K-366K Annually
Expert/Leader
244K-366K Annually
Expert/Leader
Cloud • Healthtech • Social Impact • Software • Biotech
Leads Benchling’s enterprise applications function, owning the SaaS portfolio, architecture, integrations, platform lifecycle, and vendor strategy. Oversees revenue and finance systems programs, including quote-to-cash, pricing, packaging, and billing. Drives automation, AI adoption, self-service, audit-ready SOX controls, and measurable operational efficiency. Partners with Finance, Sales, RevOps, People, Legal, and CX leaders while developing the applications team and contractor bench.
Top Skills: Artificial IntelligenceAtlassianBoomiCertiniaGainsightMulesoftNetSuiteSalesforceSalesforce CpqSalesforce Data CloudSalesforce Revenue CloudSlackSnowflakeTrayWorkatoWorkdayZendesk
2 Minutes Ago
Hybrid
Santa Clara, CA, USA
184K-288K Annually
Expert/Leader
184K-288K Annually
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads technical accounting matters, purchase accounting, M&A financial diligence, SEC reporting, accounting policy implementation, audit resolution, and quarterly and year-end close activities. Oversees complex US GAAP research, internal controls, process improvements, and cross-functional collaboration with Finance, Legal, Corporate Development, and external auditors. Develops the technical accounting team and supports AI-enabled workflow optimization.
Top Skills: Ai Accounting Automation ToolsExcelSAPSec Reporting StandardsSox 404Us Gaap
2 Minutes Ago
Hybrid
San Mateo, CA, USA
106K-120K Annually
Senior level
106K-120K Annually
Senior level
Cloud • Fintech • Information Technology • Machine Learning • Software
Drive net-new business and revenue growth across a territory of accounting and bookkeeping partners. Develop migration strategies, manage partner relationships, conduct client meetings and events, deliver certification training, and educate practices on digital accounting and the Xero platform. Maintain Salesforce data, manage sales cycles, forecast accurately, and independently execute territory plans while traveling regularly for face-to-face client engagements.
Top Skills: SalesforceXero Platform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account