STN Inc Logo

STN Inc

Hardware Engineer

Posted 2 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Owns the hardware lifecycle for GPU and infrastructure assets, including fleet health monitoring, vendor RMA workflows, firmware and BIOS upgrades, burn-in testing, failure investigation, inventory accuracy, capacity planning, spare-parts strategy, runbook creation, and new platform qualification. The role serves as the technical owner of physical compute platforms and supports high-density GPU infrastructure across sites.
The summary above was generated by AI
Hardware Engineer

Infrastructure operations · shared across sites

Reports to: Director, Hardware Engineering

Location: Pleasanton, CA (hybrid) or assigned site; travel up to 25%

Department: Infrastructure & DC Operations / Systems Engineering

Position summary

The Hardware Engineer owns hardware lifecycle for GPU and supporting infrastructure assets, including fleet health monitoring, RMA workflows, firmware management, and long-range capacity planning. The role is the technical owner of the physical compute platform.

Key responsibilities
  • Monitor GPU and server health including thermal, error rates, and component failures

  • Drive the RMA process with vendors (NVIDIA, Supermicro, HPE, and others) end-to-end

  • Manage firmware, BIOS, and BMC upgrade campaigns across the fleet

  • Develop hardware burn-in and acceptance test procedures, including NCCL and stress tests

  • Investigate hardware failures and produce vendor-grade root cause analyses

  • Maintain hardware inventory, asset records, and CMDB accuracy

  • Drive capacity planning across compute, storage, and networking

  • Coordinate with Procurement on spare parts strategy and stocking levels

  • Author hardware engineering runbooks and operational procedures

  • Support new platform bring-up, qualification, and reference architecture validation

Required qualifications
  • 5+ years in hardware engineering, systems engineering, or data center engineering

  • Deep knowledge of x86 server architecture, GPU systems, and modern storage

  • Hands-on experience with NVIDIA HGX, DGX, or hyperscale-class systems

  • Strong Linux fundamentals and scripting skills (Python, Bash)

  • Bachelor's degree in computer science, electrical engineering, or related field

Preferred qualifications
  • Experience with NVIDIA Mission Control, Base Command Manager, or Bright Cluster Manager

  • Familiarity with IPMI, Redfish, and vendor management interfaces

  • Knowledge of liquid cooling and high-density power architectures

  • Experience operating fleets of 1,000+ GPUs

Similar Jobs

Yesterday
In-Office or Remote
2 Locations
124K-226K Annually
Senior level
124K-226K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Designs and validates server subsystems and hardware platforms, ensuring functional specifications, safety, signal integrity, power integrity, and manufacturing requirements are met. Responsibilities include PCB and stackup development, board file analysis, safety compliance, hardware validation, technical documentation, and cross-functional collaboration. The role supports product development from design through qualification and production, using tools such as HFSS, Maxwell, Cadence Concept, and Allegro, along with laboratory validation equipment.
Top Skills: Cadence AllegroCadence ConceptHfssLogic AnalyzersMaxwellOscilloscopesPcbPdn AnalysisPower Integrity (Pi)Signal Integrity (Si)Spectrum AnalyzersStackup
15 Days Ago
Remote or Hybrid
USA
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Validate server hardware components and new platform designs through power, thermal, performance, firmware, and automation testing. Maintain hardware test infrastructure, document technical findings, coordinate with ODM/OEM partners, and collaborate with software teams on BIOS, kernel, and firmware optimization. The role requires expertise in server architectures, validation methodologies, Linux, Python, Bash, Redfish, IPMI, and industry standards including PCIe, NVMe, and DDR.
Top Skills: BashBiosBmcCloud InfrastructureContainerizationDdrFirmwareGpusIpmiKernelLinuxNvmeOcpOpenbmcPcie Gen6PythonRedfish
23 Days Ago
Remote or Hybrid
Senior level
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Lead development of high-fidelity component plant models in MATLAB/Simulink for electrical, mechanical, and software interfaces. Collaborate with suppliers and cross-functional teams to parameterize, validate, and integrate models into CoSim/SIL/HIL environments (FMU export). Define modeling standards, lead correlation/validation strategies, document assumptions, and mentor peers to scale virtual hardware modeling across programs.
Top Skills: CosimFmuHilMatlabSilSimulink

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account