Sciforium Logo

Sciforium

GPU Cluster Engineer, Hardware Operations

Posted One Month Ago
Be an Early Applicant
In-Office
2 Locations
150K-180K Annually
Mid level
In-Office
2 Locations
150K-180K Annually
Mid level
Own physical health and infrastructure of GPU clusters: respond to hardware outages, monitor GPU thermals/power, coordinate vendor RMAs and data center work, rack and bring up nodes, maintain Linux OS across bare-metal servers, secure networking and SSH, and manage directory and network storage services to keep compute environment stable for research and product teams.
The summary above was generated by AI

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

Role Overview

We are looking for a dedicated Hardware Operations Engineer to own the physical health and foundational infrastructure of our GPU clusters. You will be the primary custodian of our compute hardware, responsible for everything from data center vendor coordination up to the base Linux OS layer. You will ensure our research and product teams have a stable, secure, and fully operational physical environment to run their demanding compute workloads on.

Key Responsibilities

  • System Health & Hardware Reliability

    • On-Call Response: Serve as the primary point of contact for physical system outages, hardware failures, and network interruptions to minimize downtime.

    • Cluster Monitoring: Proactively monitor hardware health, including GPU thermals, power draw, and physical system loads, catching anomalies before they impact active workloads.

    • Vendor Liaison: Work closely with data center facility staff and third-party hardware vendors to coordinate RMA processes, physical repairs, part replacements, and routine maintenance.

    • Hardware Deployment: Rack, cable, and lead the physical bring-up of new GPU nodes, ensuring power and network connectivity are fully integrated into the existing cluster.

  • Linux & Network Administration

    • OS Management: Install, patch, and maintain Linux operating systems (Ubuntu/CentOS/RHEL) across the cluster bare-metal servers.

    • Security & Access: Configure and maintain edge and internal networking, including firewalls, VPNs, and strict SSH access controls to secure our infrastructure.

    • Identity & Storage Management: Administer LDAP/Active Directory for centralized user authentication and ensure network storage systems (NFS/GPFS/Lustre) are reliably mounted and properly permissioned.

Qualifications

  • Must-Haves:

    • 3+ years of experience in Linux Systems Administration (deep knowledge of boot processes, systemd, disk management, etc.).

    • Strong background in server hardware troubleshooting, specifically within high-density environments (power, cooling, PCIe topologies).

    • Experience managing networking security (VPNs, iptables/firewalld, VLANs) and directory services (LDAP/FreeIPA/Active Directory).

    • Proficiency in Bash scripting for essential system automation.

  • Nice-to-Haves:

    • Experience using configuration management tools like Ansible, SaltStack, or Terraform for OS provisioning.

    • Familiarity with data center operations, cooling requirements for high-TDP accelerators (like NVIDIA B200 or AMD MI355x).

Benefits include
  • Medical, dental, and vision insurance

  • 401k plan

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

Equal opportunity

Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

HQ

Sciforium San Francisco, California, USA Office

San Francisco, CA, United States

Sciforium Los Altos, California, USA Office

4401 El Camino Real, Los Altos, California, United States, 94022

Similar Jobs

An Hour Ago
Remote or Hybrid
United States
145K-196K Annually
Senior level
145K-196K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Manage and mentor Product Managers while guiding product strategy, prioritization, roadmap execution, and lifecycle management for Rapid7’s cybersecurity platform. Partner with Product, Engineering, UX, executives, customers, and cloud ecosystem partners to define requirements, launch features, analyze usage and feedback, and drive market impact. The role also promotes AI-enabled product management practices, customer-centric decision-making, cross-functional alignment, and development of high-performing global teams.
Top Skills: Artificial IntelligenceCloud SecurityVulnerability Management
An Hour Ago
Remote or Hybrid
United States
23-30 Hourly
Mid level
23-30 Hourly
Mid level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Provides independent healthcare consulting services for skilled nursing, independent living, and assisted living clients. Analyzes and reconciles claims, reimbursements, and operational data; prepares reports; supports client onboarding and regulatory readiness; responds to healthcare-related inquiries; and collaborates with compliance, IT, and finance teams. Requires knowledge of healthcare regulations, payer requirements, healthcare systems, and strong analytical and communication skills.
Top Skills: Claims Processing PlatformsEhrsHealthcare Analytics ToolsHipaaMacraExcelMicrosoft OutlookMicrosoft Word
3 Hours Ago
Remote or Hybrid
USA
100K-223K Annually
Senior level
100K-223K Annually
Senior level
Machine Learning • Payments • Security • Software • Financial Services
Leads strategy, roadmap, prioritization, development, and financial performance for business technology products supporting Vendor Finance. Coordinates technology, business, risk, compliance, and vendor stakeholders; manages product initiatives, timelines, budgets, customer experience, and regulatory requirements. Evaluates product opportunities, drives innovation using emerging technologies including AI, and leads product development teams, including staffing, performance management, and talent development.
Top Skills: Agile MethodologyArtificial IntelligenceSoftware Development Life Cycle

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account