Fluidstack Logo

Fluidstack

Customer Reliability Engineer

Posted 21 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
173K-224K Annually
Mid level
In-Office
San Francisco, CA, USA
173K-224K Annually
Mid level
Own reliability for named customer workloads; debug distributed systems across hardware, fabric, and scheduler; run customer-facing incident communications; convert recurring customer pain into engineering fixes; support large-scale compute customers and push internal teams to resolve root causes.
The summary above was generated by AI
About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.


We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

The Data Center Operations Team

Examples of key problems the team is working on

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.

  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.

  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.

Role Scope
  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.

  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.

  • Run customer-facing incident communication with technical depth and no spin.

  • Turn recurring customer pain into engineering fixes with the production teams.

What We're Looking For
  • The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.

  • You debug distributed systems methodically across layers you don't own.

  • You've written incident updates customers trusted more after reading.

  • You push internal teams to fix causes, not symptoms, and follow up until they do.

  • Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Similar Jobs

5 Days Ago
In-Office
Mid level
Mid level
Aerospace • Security • Energy • Defense
Works on-site at customer facilities to evaluate, select, install, commission, and monitor mechanical seals and support systems. Collects and analyzes reliability data, recommends MTBR improvements, provides training, and communicates findings using reliability software (e.g., ServiceMax). Supports account/regional engineers and may lead customer reliability personnel.
Top Skills: MS OfficeServicemax
12 Days Ago
In-Office
Internship
Internship
Aerospace • Security • Energy • Defense
Part-time Mechanical Engineering intern rotating through inspection/repair, maintenance, assembly, and engineering/sales. Inspect, disassemble, lap, assemble, pressure-test mechanical seals, perform failure/root-cause analysis, assist sales, run seal calculations with engineering, and communicate with stakeholders.
Top Skills: CadExcelMicrosoft PowerpointMicrosoft WordSolid Modeling
18 Days Ago
In-Office or Remote
125K-130K Annually
Senior level
125K-130K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Software • Analytics • Infrastructure as a Service (IaaS) • Big Data Analytics
Operate, monitor, and maintain Astronomer's managed Airflow platform and underlying cloud/Kubernetes infrastructure. Troubleshoot customer environments, participate in on-call rotation, build monitoring/automation, improve observability, and work directly with customers to meet SLAs and drive reliability.
Top Skills: Apache AirflowAWSAzureCi/CdGCPIacKubernetesLinuxPython

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account