Fluidstack Logo

Fluidstack

Distributed Systems Engineer

Reposted 18 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
208K-269K Annually
Senior level
In-Office
San Francisco, CA, USA
208K-269K Annually
Senior level
Design, build, and operate the observability and control-plane for a hyperscale GPU fleet. Own data pipelines, API surface, fleet state as source-of-truth, distributed command execution, and clean hardware/site onboarding (ZTP/DHCP/DNS). Ensure production reliability, run incidents, and drive automation to eliminate toil.
The summary above was generated by AI
About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.


We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward. ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

     

The Production Engineering Team

Examples of key exciting problems the team is working on

  • Make tens of thousands of GPUs legible in real time: build the observability platform that turns raw telemetry into signal, from site-level health down to individual device and link. At this scale, you cannot operate what you cannot see.

  • Build the control plane every team at Fluidstack depends on: replace one-off tooling with a stable, versioned API surface that covers unified machine management, actual state inspection, and distributed command execution. One interface for the whole company, not a hundred scripts.

  • Make the system's view of itself always match reality: integrate fleet state as a machine-readable source of truth across provisioning, operations, and customer-facing platforms, so every new site and GPU generation lands cleanly from day zero.

Role Scope
  • Own the observability platform. Build and operate the data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible — from site down to device and link. No other team should need to scrape production directly to answer a question.

  • Define and build the API surface for infrastructure. Design the contracts between production infrastructure and every tool that touches it. All other teams at Fluidstack use your tooling to manage and operate our hyperscale fleet.

  • Build the production control plane. Unified machine management, actual state inspection, distributed command execution — and the Kubernetes-based infrastructure that underpins it all.

  • Own fleet state as source of truth. SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms. What the system says about itself should match reality, and you're accountable when it doesn't.

  • Land new hardware into the platform cleanly. ZTP, DHCP, DNS, artifacts — every new XPU generation and site integration goes through IaaS before production.

     

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You treat toil as a bug. If something requires a human to do it twice, you build the thing that makes it not require a human.

  • You design APIs that age well. You've felt the pain of a leaky abstraction at scale and you don't repeat it.

  • You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.

  • You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.

  • You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.

  • You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.

  • You've shipped production services that other teams depend on at scale, and you're comfortable in any language using AI coding tools.

  • Bonus: Distributed systems and data pipeline engineering. Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics). API design and versioning at scale. Workflow and orchestration engines (Temporal, Cadence). BMC/Redfish or hardware telemetry. Go, Python, and Postgres.

 

Salary & Benefits
  • Competitive total compensation package (salary + equity).

  • Retirement or pension plan, in line with local norms.

  • Health, dental, and vision insurance.

  • Generous PTO policy, in line with local norms.

     

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Similar Jobs

13 Days Ago
Hybrid
Senior level
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Design, build, and operate backend services that enforce regional data residency across a globally distributed edge. Own end-to-end features from design and implementation to rollout and production operations, collaborate across platform, cryptography, and product teams, reason about failure modes and compliance tradeoffs, participate in on-call incident response, and raise engineering standards through reviews and mentorship.
Top Skills: C/C++ClickhouseCloudflare WorkersDurable ObjectsEdge/CdnEvent Streaming/Asynchronous MessagingGlobally Distributed Key-Value StorageGoGrpcHsmKubernetesL4/L7 ProxiesMetrics/Logs/TracingPkiPostgresRestRust
5 Days Ago
In-Office or Remote
2 Locations
224K-431K Annually
Senior level
224K-431K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, deploy, and run large-scale infrastructure services on NVIDIA hardware. Define SLOs and observability, automate toil, participate in on-call incident response, and consult with peer teams on systems design and reliability.
Top Skills: BmaasContainersDockerEdaGoKubernetesLinuxNcclNetworkingNvidia HardwareOpenstackPerlPythonRubySlurmStorage
4 Days Ago
Hybrid
San Francisco, CA, USA
200K-288K Annually
Expert/Leader
200K-288K Annually
Expert/Leader
Cloud • Information Technology • Security • Software • Cybersecurity
Lead technical ownership of a global platform for safe change: distributed key-value storage, progressive configuration delivery, testing and health-mediation. Write production code, design systems, review senior proposals, integrate LLMs, and drive cross-team initiatives to improve reliability, rollout safety, and developer productivity.
Top Skills: Api DesignGoogle AnnealingGoogle ProdspecKey-Value StoresKubernetesLlmsRocksdb

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account