ifm Logo

ifm

HPC Engineer

Reposted 28 Days Ago
Be an Early Applicant
In-Office
Sunnyvale, CA, USA
150K-300K Annually
Junior
In-Office
Sunnyvale, CA, USA
150K-300K Annually
Junior
Provide overnight operational coverage for large-scale GPU clusters: monitor health and performance, triage incidents, support researchers, run recovery procedures, validate deployments and upgrades, track utilization, and build automation and monitoring tools.
The summary above was generated by AI
About MBZUAI
The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.
Position Summary
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.

Responsibilities

    • Monitor health, performance, and availability of large-scale GPU clusters.
    • Respond to incidents and perform first-level triage.
    • Support researchers and troubleshoot job failures.
    • Execute operational runbooks and recovery procedures.
    • Validate cluster deployments, upgrades, and maintenance activities.
    • Track infrastructure utilization and operational metrics.
    • Develop automation and monitoring tools.
    • Contribute to documentation and reporting.

Education

    Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

Experience

    • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
    • Strong Linux troubleshooting skills.
    • Experience with scripting using Python or Bash.

Preferred Qualifications

    • Slurm.
    • GPU infrastructure.
    • AWS, Azure, or GCP.
    • Grafana, Prometheus, Datadog, or similar tools.
    • Containers and Kubernetes.
    • AI/ML infrastructure exposure.
    • Research computing environments.

Benefits Include
*Comprehensive medical, dental, and vision benefits 
 *Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability
 

Similar Jobs

5 Hours Ago
In-Office
Milpitas, CA, USA
136K-200K Annually
Senior level
136K-200K Annually
Senior level
Hardware
Develop and optimize high-performance computational systems for semiconductor product development. Build reusable Python modules, parallel computing models, and workload benchmarks across Linux, GPU, server, storage, and interconnected network environments. Configure hardware, operating systems, high-availability infrastructure, observability exporters, and configuration management tools. Support electrical engineering standards, critical facility infrastructure, renewable energy initiatives, and cross-functional technical strategy while translating requirements into design guidance and executive-ready summaries.
Top Skills: AlmalinuxAnsibleBattery StorageBiosChefCudaDockerElasticsearchEthernetEv ChargingGitGpusGrafanaHbaInfinibandJenkinsLinuxLinux HaMdraidMpiOpenclOpenmpiPkiPrometheusPuppetPxePythonRaidRdmaRed HatRoce V2Rocky LinuxSaltSolarSsl/TlsSuseSystemdUbuntuUpsX86
6 Days Ago
In-Office
Milpitas, CA, USA
136K-232K Annually
Senior level
136K-232K Annually
Senior level
Hardware
Develop and troubleshoot large-scale C++ software for multidisciplinary hardware/software systems and semiconductor products. Responsibilities include system-level programming, software architecture, multithreading, memory optimization, performance profiling, debugging, unit testing, database design, and CI/CD practices on Windows and Linux. The engineer will collaborate across staff and management levels and travel to customer sites on short notice to troubleshoot software issues.
Top Skills: Ai ToolsC++Ci/CdGitLinuxMemory ManagementMultiprocessingMultithreadingObject-Oriented DesignPerformance ProfilingRelational DatabasesSoftware ArchitectureWindows
18 Days Ago
In-Office
Sunnyvale, CA, USA
175K-200K Annually
Senior level
175K-200K Annually
Senior level
Hardware • Semiconductor • Manufacturing
Serve as the technical counterpart to the commercial team for HPC and computational engineering customers. Lead workload discovery, performance analysis, benchmarking, proofs of concept, technical evaluations, and deployment planning. Translate customer needs into engineering requirements, explain Bolt’s GPU architecture, resolve technical blockers, develop technical materials, and assess technical fit. Support strategic partners, national laboratories, government organizations, and industry events while bringing customer feedback into product strategy.
Top Skills: CC++Commercial Simulation SoftwareComputational EngineeringCpuCudaDistributed ComputingElectromagnetic SimulationFdtdFemGpuHeterogeneous ComputingHigh-Speed InterconnectsHpcHpc ClustersMomMpiNumerical SimulationPerformance OptimizationPerformance ProfilingPythonScientific Computing

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account