NVIDIA Logo

NVIDIA

Senior Solutions Architect, NVIDIA Cloud Partner Operations

Posted 3 Hours Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Santa Clara, CA, USA
224K-357K Annually
Senior level
In-Office or Remote
Hiring Remotely in Santa Clara, CA, USA
224K-357K Annually
Senior level
Lead Day 2 operations improvements across NVIDIA Cloud Partners operating GPU, HPC, and AI cloud infrastructure. Partner with engineering teams to troubleshoot production issues, improve reliability, performance, utilization, recovery, and economics, and prepare new NVIDIA technologies for operations. Raise operational maturity across tooling, telemetry, security, incident response, and automation. Convert validated solutions into procedures, architectures, assessments, and workflows while providing cross-partner feedback to NVIDIA engineering and product teams.
The summary above was generated by AI

NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem. Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time. You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.

Our job is to work hand in hand with NCPs to solve real problems and drive real optimizations, prove the answer, and turn it into something the next partner can use! This is not an outsourced operations role. The partner owns its cloud; success means leaving its team more capable, not more dependent on ours.
 

What you'll be doing:

  • Solve hard Day 2 operations problems at scale. Work alongside partner engineers to find the cause, prototype an approach, validate it under representative load, and leave behind a practice their team can operate.

  • Make new technology Day 2 ready. Help partners prepare the operating model for new NVIDIA platforms, capacity, services, and use cases before customers depend on them, and help drive adoption in live environments without degrading service.

  • Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per token to show where the cloud is losing performance or margin - and whether the fix worked.

  • Raise each partner's Day 2 maturity. Identify and help close the gaps that matter across people, process, tooling, telemetry, security, and incident response.

  • Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows that other NCPs can integrate into their standard operating model.

  • Create the feedback loop only NVIDIA can. Spot patterns across partners early and bring clear field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level.

What we need to see:

  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.

  • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.

  • Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.

  • Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and driver lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.

  • Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.

  • Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.

  • A detailed evidence-led approach to troubleshooting across system boundaries, paired with the judgment to make difficult technical findings clear.

  • The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority or taking ownership away from the operator.

  • Strong communication, prioritization, and time-management skills across multiple partner engagements.

Ways to stand out from the crowd:

  • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load.

  • Built or matured a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design.

  • Hands on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72 into production, or have hands-on experience with NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and the GPU or Network Operators.

  • Driven improved fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation.

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. This role presents an opportunity to have a wide impact at NVIDIA by improving the factory planning function. Are you creative, hard-working, dedicated, and determined? Do you love a challenge? If so, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 25, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Santa Clara, California, USA Office

2701 San Tomas Expressway, Santa Clara, CA, United States, Santa Clara

NVIDIA San Francisco, California, USA Office

San Francisco, United States

NVIDIA San Jose, California, USA Office

San Jose, United States

Similar Jobs

10 Days Ago
In-Office or Remote
Santa Clara, CA, USA
224K-357K Annually
Senior level
224K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead Day 2 operations engagements with NVIDIA Cloud Partners: diagnose and fix production GPU/HPC/cloud issues, prototype and validate fixes, improve reliability, performance, and economics, and convert solutions into reusable procedures, automation, and feedback for product and engineering teams.
Top Skills: AnsibleArgo CdBase Command ManagerBashBmc/RedfishDcgmFirmwareGb200Gb300 Nvl72GitopsGpu OperatorGpu SchedulingGrafanaHigh-Speed EthernetIbm Storage ScaleInfinibandKubernetesLinuxLustreMission ControlNcclNetwork OperatorOpentelemetryPrometheusPythonSlurmSpectrum-XTerraformUfmVast DataWeka
42 Minutes Ago
Remote or Hybrid
USA
65K-138K Annually
Junior
65K-138K Annually
Junior
Consumer Web • Coupons • Healthtech • Social Impact • Pharmaceutical
Supports broad-reach paid media and integrated marketing campaigns across television, streaming video and audio, radio, and podcasts. Responsibilities include analyzing channel performance, recommending optimizations, monitoring digital marketing results, managing budgets and pacing, handling invoicing, building tracking assets, supporting campaign planning and launches, conducting competitive analysis, and presenting performance insights. The role collaborates with internal teams and external agency partners.
Top Skills: C-TrackingCtvGoogle SuiteLinear TvMS OfficeOlvOttPaid MediaPodcastsQr CodesStreaming AudioStreaming VideoTerrestrial RadioUtm Tracking
59 Minutes Ago
Remote or Hybrid
United States
17-25 Hourly
Junior
17-25 Hourly
Junior
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Provides routine product-use assistance and technical support for Dealertrack products. Troubleshoots and documents system issues, logs customer cases in CRM, manages multiple tickets, meets service-level targets, follows up with customers, and coordinates with internal departments. Requires flexible availability for business-hour shifts, including Saturdays and overtime, along with strong communication, problem-solving, independence, and teamwork.
Top Skills: CRMDealertrack DmsGenesys CloudSalesforce

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account