Spellbrush Logo

Spellbrush

HPC/ML Infrastructure Engineer

Reposted One Month Ago
In-Office
San Francisco, CA, USA
Mid level
In-Office
San Francisco, CA, USA
Mid level
Lead bring-up, administration, and operations of a large GPU/AI training cluster. Serve as bridge between researchers and hardware, ensuring SLURM jobs, parallel filesystems, networking, and monitoring operate reliably. Work across provisioning, storage, VPN/access, and traditional Linux sysadmin tasks; assist with physical racking and on-site datacenter needs. Collaborate closely with a small research team in Tokyo or San Francisco.
The summary above was generated by AI

We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.

You may be a good fit if:You love anime and the anime aesthetic.

This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.

You’re familiar with the modern HPC software landscape

Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)

As well as the traditional sysadmin landscape

Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools.

You're not afraid of physical computers

We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.

And you're comfortable working on small, fast-paced teams.

We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world.

We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.

HQ

Spellbrush San Francisco, California, USA Office

San Francisco, California, United States

Similar Jobs

Entry level
Artificial Intelligence • Cloud • Information Technology • Consulting
Provides technical service and post-sales support for customer environments. Resolves routine incidents, collaborates on complex issues, applies HPE solutions, manages customer expectations, and builds account relationships. The role requires developing technical knowledge across areas such as server administration, security, performance management, and IT environments while delivering customer-focused solutions within defined parameters.
Top Skills: Customer Relationship Management (Crm)Infrastructure As A Service (Iaas)Performance ManagementServer AdministrationTechnical Security Management
2 Days Ago
In-Office
Mid level
Mid level
Artificial Intelligence • Cloud • Information Technology • Consulting
Provides onsite post-sales technical support for customer server, networking, and business systems environments. Resolves routine incidents, collaborates on complex issues, applies HPE solutions, and aligns technical services with customer needs. The role includes proactive and reactive service delivery under SLAs, customer expectation management, account relationship building, and communication with technical stakeholders and first-level management.
Top Skills: Clustered Server EnvironmentsInfrastructure As A Service (Iaas)NetworkingPerformance ManagementServer AdministrationTechnical Security Management
2 Days Ago
Hybrid
Entry level
Entry level
Artificial Intelligence • Cloud • Information Technology • Consulting
Provide consulting and technical support for HPC facility infrastructure, including electrical systems, power equipment, air conditioning, and water-cooling systems. Responsibilities span proposal development, facility and equipment design, construction and installation support, commissioning, maintenance, and operational support. The role coordinates with customers, facility personnel, internal teams, and contractors in Japanese and English, requiring knowledge of electrical and cooling systems, construction management, and CAD drafting.
Top Skills: AutocadElectrical SystemsHpcJw CadMechanical Cooling SystemsExcelMicrosoft PowerpointMicrosoft WordWater Cooling Systems

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account