NVIDIA Logo

NVIDIA

Software Platform Support Engineer - GPU Cloud

Posted 3 Days Ago
Be an Early Applicant
In-Office or Remote
5 Locations
108K-173K Annually
Senior level
In-Office or Remote
5 Locations
108K-173K Annually
Senior level
Provide Tier 1 support for complex cloud platforms, troubleshoot distributed software and customer issues, investigate root causes, improve operational workflows, create runbooks and documentation, build support tooling, and coordinate with engineering, SRE, and other internal teams. The role supports production systems through an on-call rotation and requires expertise across cloud infrastructure, networking, storage, Kubernetes, Linux, and DevOps tooling, with GPU, MLOps, HPC, or SLURM experience preferred.
The summary above was generated by AI

The NVIDIA DGX Cloud organization is looking for passionate software support engineers to partner closely with our internal customers to support them on our internal platforms. This partnership requires you to gain a deep understanding of the customer needs, how their application(s) work, assist them in troubleshooting issues, and create documentation to make it easier for users to troubleshoot issues themselves in an ambiguous / fast-moving environment. The support you provide will help our users have a better experience and help shape our platform. 

 

We expect you to have knowledge of supporting cloud-based deployments across compute, storage and networking environments. 

 

What will you be doing:

  • Coordinate with multiple internal teams to provide Tier 1 support for complex cloud platforms

  • Define and improve operational workflows (runbooks, escalation paths, support processes)

  • Triage/investigate root cause of customer issues and escalate as needed 

  • File bugs and report issues while working closely with the Site Reliability team

  • Build tooling to improve customer support process and visibility

  • Deeply understand user workloads and use cases 

  • Partner with multiple internal teams to give feedback to engineering teams and develop solutions to aid in their success

  • Be part of an on call rotation to support production systems

 

What we need to see:

  • BS/MS degree in Computer science or related areas (or equivalent experience)

  • 5+ yrs of experience with supporting distributed software systems, supporting end-user software platforms, and experience with Linux

  • Experience with Kubernetes, AWS, Azure, OCI, and GCP 

  • Background of Infrastructure, Networking, Storage, and DevOps scripting/tooling

  • Understanding of data storage technologies (databases, file, block, blob)

  • Customer Service/Support Experience

  • Willingness to work up and down the stack as well as across multiple teams 

  • Strong skills in troubleshooting and Communication

 

Ways to stand out from the crowd:

  • Experience with MLOps workflows or ML infrastructure 

  • Familiarity with GPU workloads or distributed training systems

  • SLURM or HPC previous experience

  • Strong drive to work with internal customers and make them successful

  • A drive to improve process with strong organizational skills  

 

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 108,000 USD - 172,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 15, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Santa Clara, California, USA Office

2701 San Tomas Expressway, Santa Clara, CA, United States, Santa Clara

NVIDIA San Francisco, California, USA Office

San Francisco, United States

NVIDIA San Jose, California, USA Office

San Jose, United States

Similar Jobs

27 Minutes Ago
In-Office or Remote
San Francisco, CA, USA
200K-260K Annually
Expert/Leader
200K-260K Annually
Expert/Leader
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Lead regional growth strategy for USDC by building partnerships with exchanges and OTC desks, delivering regionally compliant product initiatives, and creating AI-driven, automated growth platforms. Own experimentation frameworks, dashboards, and metrics to scale liquidity, adoption, and measurable share shift versus competing stablecoins through cross-functional execution and data-driven prioritization.
Top Skills: Agent-Driven SystemsAIAnalyticsAutomation PlatformsBlockchainDashboardsExperimentation FrameworksStablecoins
27 Minutes Ago
In-Office or Remote
San Francisco, CA, USA
200K-258K Annually
Expert/Leader
200K-258K Annually
Expert/Leader
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Leads Circle’s internal communications and employer brand strategy across a global, distributed workforce. Partners with executives, HR, Talent Acquisition, Marketing, and Corporate Communications to shape company messaging, employee engagement, EVP, talent storytelling, candidate experience, alumni programs, and employer reputation. Oversees content, communication infrastructure, change communications, measurement, and AI-enabled workflows while managing and developing a communications team.
Top Skills: AILlm
55 Minutes Ago
Remote or Hybrid
5 Locations
123K-123K Annually
Expert/Leader
123K-123K Annually
Expert/Leader
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead and oversee multiple multidisciplinary teams to architect, develop, deploy, monitor, and maintain full-stack AI solutions (including LLM integrations). Provide strategic and technical leadership, drive experimentation and CI/CD practices, ensure quality and stakeholder alignment, promote innovation, and mentor senior and junior team members to deliver business-aligned AI outcomes.
Top Skills: AWSAzureCi/CdGitGoogle Cloud PlatformLangchainLlmsPandasPrompt EngineeringPythonPyTorchScikit-LearnSemantic KernelSQLVector Dbs

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account