NVIDIA Logo

NVIDIA

Senior Engineer System Software, SDN Operations

Posted 2 Days Ago
Be an Early Applicant
In-Office or Remote
2 Locations
184K-288K Annually
Senior level
In-Office or Remote
2 Locations
184K-288K Annually
Senior level
Design, develop, and operate large-scale SDN control and data plane software (OVS/OVN/OpenFlow) for NVIDIA AI Cloud. Build network orchestration and IaC (gRPC/REST, Ansible, Terraform, ArgoCD/Flux), maintain CI/CD pipelines, implement observability/monitoring, drive upstream open-source contributions, and collaborate with SRE/DevOps to ensure production reliability and performance SLAs.
The summary above was generated by AI

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions.

This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response.

What you'll be doing:

  • Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow)

  • Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes

  • Drive upstream contributions to OVN-Kubernetes and related open-source projects

  • Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis

  • Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments

  • Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs

  • Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs

  • Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure

  • Drive reliability through incident management, resource monitoring, and performance tuning

  • Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational tooling

What we need to see:

  • BS/MS in Computer Science or related technical field, or equivalent experience

  • 8+ years of proven experience in software development for large-scale distributed environments

  • Expert-level knowledge of OVN, OVS, OpenFlow, and modern network protocols

  • Strong programming skills in C and Go; advanced scripting in Bash and Python

  • Deep knowledge of Kubernetes, practical experience deploying and supporting CNIs (OVN-Kubernetes)

  • Hands-on experience with Infrastructure-as-Code and deployment tools (Ansible, Terraform, ArgoCD, Flux)

  • Experience designing and operating complex, multi-stage CI/CD pipelines

  • Hands-on experience developing secure, high-performance services using gRPC and REST with TLS and strong authentication

  • Strong knowledge of datacenter routing, switching, and Linux host/VM networking

Ways to stand out from the crowd:

  • Contributions to open-source projects (especially OVS, OVN, OVN-Kubernetes, or other Kubernetes networking projects)

  • Experience with hardware acceleration (GPU, DPU or equivalent experience) for networking 

  • Practical experience with major cloud providers (AWS, Azure, GCP) and hybrid/multi-cloud deployments

  • SRE/DevOps top-level expertise — on-call, incident management, operations focused on service reliability targets, production ownership

  • Experience with observability platforms and tools (Prometheus, Grafana, Jaeger, OpenTelemetry, ELK)

With a competitive salary package and benefits, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Are you a creative and autonomous Systems Software Engineer who loves challenges? Do you have a genuine passion for advancing the state of Networking, SDN, and Cloud Networking across a variety of industries? If so, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 18, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Santa Clara, California, USA Office

2701 San Tomas Expressway, Santa Clara, CA, United States, Santa Clara

NVIDIA San Francisco, California, USA Office

San Francisco, United States

NVIDIA San Jose, California, USA Office

San Jose, United States

Similar Jobs

6 Hours Ago
Remote or Hybrid
Palo Alto, CA, USA
100K-176K Annually
Mid level
100K-176K Annually
Mid level
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Design, implement, and operate critical, scalable backend services (identity, friend graph, storage). Collaborate across teams, evaluate trade-offs, test and debug, manage availability, scalability, operational excellence, cost, and participate in incident resolution.
Top Skills: AWSC++GCPJavaKotlinKubernetesMemcacheNoSQLPythonRedisSwift
13 Hours Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
The role involves designing cloud infrastructure, managing production Kubernetes clusters, optimizing CI/CD pipelines, enhancing developer experience, and ensuring reliable AI workloads. Candidates should have extensive experience in infrastructure and distributed systems engineering with strong coding skills and cloud expertise.
Top Skills: AWSAzureDatadogDockerElkGCPGoGrafanaJavaKubernetesPrometheusPythonTerraform
18 Hours Ago
Remote or Hybrid
United States
17-25 Hourly
Junior
17-25 Hourly
Junior
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Provide remote technical support for VinSolutions and Cox Automotive products via phone, email, and chat. Troubleshoot and resolve application issues, log cases in the CRM, escalate to other teams as needed, and keep clients informed while meeting quality standards and shift requirements.
Top Skills: Genesys PurecloudExcelMicrosoft OutlookMicrosoft WordSalesforceVinsolutions

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account