Sciforium Logo

Sciforium

GPU Cluster Engineer, Hardware Operations

Posted 7 Days Ago
Be an Early Applicant
In-Office
2 Locations
150K-180K Annually
Mid level
In-Office
2 Locations
150K-180K Annually
Mid level
Own physical health and infrastructure of GPU clusters: respond to hardware outages, monitor GPU thermals/power, coordinate vendor RMAs and data center work, rack and bring up nodes, maintain Linux OS across bare-metal servers, secure networking and SSH, and manage directory and network storage services to keep compute environment stable for research and product teams.
The summary above was generated by AI

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

Role Overview

We are looking for a dedicated Hardware Operations Engineer to own the physical health and foundational infrastructure of our GPU clusters. You will be the primary custodian of our compute hardware, responsible for everything from data center vendor coordination up to the base Linux OS layer. You will ensure our research and product teams have a stable, secure, and fully operational physical environment to run their demanding compute workloads on.

Key Responsibilities

  • System Health & Hardware Reliability

    • On-Call Response: Serve as the primary point of contact for physical system outages, hardware failures, and network interruptions to minimize downtime.

    • Cluster Monitoring: Proactively monitor hardware health, including GPU thermals, power draw, and physical system loads, catching anomalies before they impact active workloads.

    • Vendor Liaison: Work closely with data center facility staff and third-party hardware vendors to coordinate RMA processes, physical repairs, part replacements, and routine maintenance.

    • Hardware Deployment: Rack, cable, and lead the physical bring-up of new GPU nodes, ensuring power and network connectivity are fully integrated into the existing cluster.

  • Linux & Network Administration

    • OS Management: Install, patch, and maintain Linux operating systems (Ubuntu/CentOS/RHEL) across the cluster bare-metal servers.

    • Security & Access: Configure and maintain edge and internal networking, including firewalls, VPNs, and strict SSH access controls to secure our infrastructure.

    • Identity & Storage Management: Administer LDAP/Active Directory for centralized user authentication and ensure network storage systems (NFS/GPFS/Lustre) are reliably mounted and properly permissioned.

Qualifications

  • Must-Haves:

    • 3+ years of experience in Linux Systems Administration (deep knowledge of boot processes, systemd, disk management, etc.).

    • Strong background in server hardware troubleshooting, specifically within high-density environments (power, cooling, PCIe topologies).

    • Experience managing networking security (VPNs, iptables/firewalld, VLANs) and directory services (LDAP/FreeIPA/Active Directory).

    • Proficiency in Bash scripting for essential system automation.

  • Nice-to-Haves:

    • Experience using configuration management tools like Ansible, SaltStack, or Terraform for OS provisioning.

    • Familiarity with data center operations, cooling requirements for high-TDP accelerators (like NVIDIA B200 or AMD MI355x).

Benefits include
  • Medical, dental, and vision insurance

  • 401k plan

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

Equal opportunity

Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

HQ

Sciforium San Francisco, California, USA Office

San Francisco, CA, United States

Sciforium Los Altos, California, USA Office

4401 El Camino Real, Los Altos, California, United States, 94022

Similar Jobs

23 Minutes Ago
In-Office
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead design and implementation of secure-by-default infrastructure: build IaC modules, CI/CD templates, policy-as-code guardrails, and platform security tooling. Embed pipeline and supply-chain security, enforce least-privilege identity, operate defenses as self-service products, review high-risk designs, and set technical direction while mentoring engineers.
Top Skills: Admission ControllersArtifact IntegrityAzure Entra IdAzure Management GroupsAzure PolicyBashCi/CdConftestContainer/Image Scanning And HardeningCosignDependency ScanningGatekeeperGoKubernetesOpaOrg PolicyPowershellPre-CommitPythonRustSbomsSecrets ManagementService Control Policies (Scps)SigstoreSlsaTerraform
24 Minutes Ago
Remote or Hybrid
140K-245K Annually
Senior level
140K-245K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead paid media measurement strategy to connect campaign performance to pipeline and buying-group engagement. Define KPIs, govern taxonomy, synthesize multi-source data, interpret attribution and MMM results, and translate insights into budget, channel, and audience recommendations. Partner with media agencies, data science, and stakeholders to improve measurement maturity, run Test & Learn programs, and present executive-level data storytelling driving business impact.
Top Skills: Adtech PlatformsAIMedia Mix Modeling (Mmm)Multi-Touch AttributionPythonSQL
24 Minutes Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead adoption and implementation of the Armis platform for platinum customers, delivering hands-on technical configuration, aligning use cases to CISO goals, building trusted relationships, coordinating cross-functional teams, advising on integrations, and driving value realization and growth opportunities.
Top Skills: ArmisIotOt/IcsPythonVeza

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account