NVIDIA Logo

NVIDIA

Senior DevOps Engineer

Posted 3 Days Ago
Be an Early Applicant
In-Office
Santa Clara, CA, USA
184K-357K Annually
Senior level
In-Office
Santa Clara, CA, USA
184K-357K Annually
Senior level
Lead the design, automation, and operation of scalable Linux infrastructure supporting networking software development and testing. Build infrastructure-as-code, configuration management, self-service tools, monitoring, and reliability practices across physical servers, networks, virtualization, containers, and storage. Diagnose complex hardware-to-application issues, establish operational standards, collaborate across teams, and mentor engineers while driving infrastructure initiatives to completion.
The summary above was generated by AI

As a Senior DevOps Engineer, you will help lead the evolution of infrastructure operations within our Networking Software group. Building on a strong Linux systems administration foundation, you will build, automate, and operate scalable platforms that support networking software development and testing. This role offers an outstanding opportunity to work with elite technology and collaborate with ambitious engineers across global sites. If you are passionate about automation, reliability, technical leadership, and continuous improvement, this is the perfect opportunity for you!

What you'll be doing:

  • Build, provision, configure, and maintain scalable Linux infrastructure for networking feature creation and validation, including physical servers, network switches, virtualization platforms, containers, and remote-management interfaces.

  • Develop automation for infrastructure provisioning, configuration management, software deployment, upgrades, and day-to-day operations using infrastructure-as-code and configuration-management practices.

  • Build reusable tools and self-service capabilities that simplify infrastructure operations, improve engineering efficiency, and reduce repetitive manual work.

  • Diagnose and resolve complex issues spanning hardware, firmware, operating systems, virtualization, containers, storage, network communications, and application environments.

  • Implement monitoring, observability, capacity management, and reliability practices to improve infrastructure performance, availability, and operational readiness.

  • Partner with engineering, IT, facilities, security, and network teams to define technical standards, maintain documentation and runbooks, and establish scalable operational processes.

  • Provide technical leadership, guide infrastructure initiatives, and mentor team members in automation, troubleshooting, and operational guidelines.

What we need to see:

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.

  • 6+ years of experience in systems engineering, DevOps, site reliability engineering, or infrastructure operations, including significant hands-on experience coordinating production or engineering Linux environments.

  • Experience working in the semiconductor industry or a hardware-focused engineering environment, with deep hands-on expertise in bare-metal Linux systems and server components—including CPUs, GPUs, memory, PCIe devices, NICs, storage, BIOS/UEFI, BMC/IPMI/Redfish, power, and cooling.

  • Proficiency in generative AI tools and skill in applying them effectively to automation, troubleshooting, documentation, operational analysis, and engineering efficiency.

  • Skilled at diagnosing complex issues across hardware, firmware, and operating-system layers.

  • Strong data-center networking knowledge, including TCP/IP, DNS, DHCP, VLANs, routing, switching, firewalls, and network troubleshooting tools.

  • Strong analytical, problem-solving, written communication, and cross-departmental collaboration skills, with the ability to guide technical initiatives to completion.

Ways to stand out from the crowd:

  • Experience managing Linux KVM/QEMU virtualization, Kubernetes clusters, multi-user engineering lab environments, NFS or distributed storage systems, and automated OS or cluster provisioning platforms.

  • Experience supporting fast-growing engineering labs, large-scale data-center environments, or globally distributed infrastructure.

  • Familiarity with observability platforms, including metrics, logging, tracing, alerting, incident management, and service-level objectives.

  • Experience crafting self-service infrastructure platforms and reusable automation that improves developer efficiency and theaAbility to establish clear, reliable, and scalable engineering and operational practices from evolving requirements.

  • Proven technical leadership, team leadership, or managerial experience, including mentoring engineers, prioritizing work, coordinating cross-functional initiatives, and driving projects from planning through completion.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 21, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Santa Clara, California, USA Office

2701 San Tomas Expressway, Santa Clara, CA, United States, Santa Clara

NVIDIA San Francisco, California, USA Office

San Francisco, United States

NVIDIA San Jose, California, USA Office

San Jose, United States

Similar Jobs

23 Days Ago
In-Office
150K-225K Annually
Senior level
150K-225K Annually
Senior level
Aerospace • Artificial Intelligence • Hardware • Machine Learning • Software • Defense • Manufacturing
The Senior DevOps Engineer will build and maintain CI/CD pipelines, Kubernetes deployments, Azure cloud infrastructure, Terraform modules, Docker images, monitoring and logging systems. The role includes strengthening infrastructure security and compliance, supporting software teams with troubleshooting and performance optimization, deploying applications in challenging customer environments, and evaluating development tools. The position focuses on spacecraft modeling and simulation software and requires hands-on design, debugging, and production infrastructure management.
Top Skills: AksAWSAzureBashCi/CdDockerGCPGithub ActionsGrafanaHelmHttp/HttpsHybrid Cloud NetworkingKubernetesLinuxLokiPrometheusPythonTcpTerraform
26 Days Ago
Hybrid
San Jose, CA, USA
186K-282K Annually
Senior level
186K-282K Annually
Senior level
Artificial Intelligence • Fintech • Software
Own FloQast’s multi-region AI platform infrastructure across AWS, including Bedrock model runtimes, sandboxed code execution, Terraform, observability, cost controls, CI/CD, security, and reliability. Support Transform, AI Matching, and AutoBuilder by managing scaling, throttling, tenant isolation, journey-based SLOs, model and prompt delivery gates, and AI-specific on-call processes across US, EU, and AU regions.
Top Skills: SparkAws AlbAws BedrockAws Bedrock AgentcoreAws EcsAws FargateAws IamAws LambdaAws NlbAws S3Aws SqsAws VpcDockerEmrFinopsGithub ActionsGrafanaHarnessMongoDBNode.jsNxOpentelemetryPostgresPrometheusPythonSnowflakeTerraformTruefoundryTypescript
Yesterday
In-Office
127K-159K Annually
Senior level
127K-159K Annually
Senior level
Aerospace
Lead DevOps practices, release management, and microservices deployments across Kubernetes, AWS, Azure, on-premises, and cloud environments. Build secure CI/CD pipelines, implement vulnerability management and observability, automate incident response, and support secure software development. Collaborate with developers and stakeholders on architecture, requirements, and DevOps processes while establishing agile teams and foundational engineering practices.
Top Skills: ArgocdAws CloudformationAws EksAzure AksBashCi/CdContainer Image ScanningDependency Vulnerability ManagementElk StackExcelGrafanaHelmKubernetesMS OfficeOpentelemetryOpentofuPowerPointPowershellPrometheusPythonSplunkTerraformWordZero Trust

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account