Fluidstack Logo

Fluidstack

Data Engineer

Reposted 7 Days Ago
In-Office
San Francisco, CA, USA
224K-279K Annually
Mid level
In-Office
San Francisco, CA, USA
224K-279K Annually
Mid level
Build and operate production data pipelines integrating ERP/ATS/project management/telemetry into a queryable layer. Own the data model for a live knowledge graph. Ship SLA-backed datasets/services for tools and ML. Convert unstructured vendor/field data (PDFs, spreadsheets) into trusted structured inputs. Ensure data quality via tests, monitoring, and lineage, and use modern data stacks and LLM tools responsibly.
The summary above was generated by AI
About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.


We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.

  • Insane urgency. We drive everything forward as fast as possible.

  • Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

  • Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.

The Ontology Team

Examples of key problems the team is working on

  • Reduce the latency from something happening to Fluidstack knowing about it. Index operational systems, documents, conversations, and telemetry as they change, so a delayed shipment or a rack coming online becomes usable context while there is still time to act.

  • Give every entity in Fluidstack's domain a shared identity. Connect a machine’s design, order, delivery, deployment, and incidents even when each system describes it differently.

  • Turn what Fluidstack knows into business signals. Compare plans with observations across sites, supply chains, and compute operations, and surface the gaps that matter as the company builds at gigawatt scale.

  • Build a secure platform where agents can be first-class citizens. Give agents access within each user’s permissions, with every answer traceable to its sources, freshness, and gaps in the evidence.

Role Scope
  • Build the internal data platform that turns updates from procurement, construction, compute operations, documents, and conversations into fresh, queryable context for Fluidstack teams and agents.

  • Make new sources straightforward to add, including vendor records, PDFs, spreadsheets, and exports, while keeping a clear path from structured data back to its original evidence.

  • Ship query and retrieval services that enforce a caller’s permissions before data reaches an agent, including when an answer combines or summarizes information from multiple systems.

  • Preserve source ownership, provenance, history, freshness, and access policies as data is transformed or combined, and propagate deletions and permission changes to downstream results.

  • Set and operate freshness, quality, and security guarantees, using real operational questions and failures to judge whether the platform helps Fluidstack act sooner.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You’ve built and operated a shared data system that changed how other teams made decisions or ran critical workflows.

  • You’ve made heterogeneous operational data coherent and extensible enough for other engineers to bring new domains into it.

  • You’ve built permissioned query or retrieval systems that software could use on behalf of people, with source restrictions enforced in search and derived answers.

  • You’ve kept data systems correct as sources, permissions, and records changed, including when data had to be replayed or removed.

  • You’ve worked directly with operators to find the decisions their data needed to support, then built shared capabilities around those needs.

  • You’ve traced a wrong or stale answer back to its source, fixed the cause, and made that failure detectable the next time.

  • Bonus: Data warehouse and lakehouse platforms (Snowflake, Databricks, ClickHouse). Change data capture and streaming (Debezium, Kafka). Authorization systems (OpenFGA, attribute-based access control). Search and document extraction. Physical infrastructure, inventory, or supply chain data.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Similar Jobs

Yesterday
Hybrid
98K-158K Annually
Entry level
98K-158K Annually
Entry level
Artificial Intelligence • Cloud • Internet of Things • Software • Cybersecurity • Industrial • Industrial Equipment
Build and maintain scalable data platforms, batch and real-time pipelines, cloud data solutions, databases, and data integration processes. Improve automation, scalability, monitoring, testing, and data quality while operationalizing data jobs and models. Partner with data scientists, analysts, and business stakeholders to deliver trusted data products supporting pricing decisions, analytics, automation, and AI-enabled outcomes.
Top Skills: Amazon RdsAmazon S3Amazon SagemakerAws BedrockAws FargateAws IamAws LambdaAzure DevopsCi/CdCodexCursorGitGithub ActionsGithub CopilotOraclePostgresPythonSnowflake
3 Days Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
99K-167K Annually
Mid level
99K-167K Annually
Mid level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Build and maintain reliable data pipelines, computed tables, and scalable data models using SparkSQL and PySpark. Integrate IoT, product, customer, external, unstructured, and ML-related data into a central data lake for analytics, causal inference, model training, and dashboards. Ensure high data quality and uptime while collaborating with data science, analytics, AI/ML, and engineering teams.
Top Skills: Apache AirflowAWSAzureDagsterData LakeData ModelingData PipelinesDatabricksDelta LakeETLGCPGitGitPrefectPysparkPythonRest ApisSparksqlSQL
7 Days Ago
Hybrid
179K-205K Annually
Mid level
179K-205K Annually
Mid level
Fintech • Machine Learning • Payments • Software • Financial Services
Design, build, and operate scalable cloud data platforms and pipelines using Python, Spark, SQL, Databricks, Snowflake, and relational or NoSQL databases. Develop distributed, real-time, and secure data solutions; enforce engineering standards; collaborate with product, engineering, analytics, and machine learning teams; communicate technical outcomes; and mentor data professionals.
Top Skills: AirflowAmazon EmrAmazon RedshiftAWSAws GlueCassandraDagsterDatabricksDynamoDBGCPJavaAzureMongoDBMonte CarloNosql DatabasesPythonScalaSnowflakeSparkSplunkSQL

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account