fal Logo

fal

Software Engineer, Distributed Systems

Reposted 14 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
180K-250K Annually
Expert/Leader
In-Office
San Francisco, CA, USA
180K-250K Annually
Expert/Leader
As a Staff Software Engineer, you will develop and maintain a core Python platform for managing computation workloads and cloud infrastructure, while ensuring system reliability and scalability.
The summary above was generated by AI

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

About this role: 

You are an experienced software engineer who thrives on building large-scale computing platforms. You have deep expertise in large scale distributed systems that deal with high complexity, a lot of traffic and data. You know how to achieve reliability and scale with minimum operational load.

Key responsibilities
  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc
  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world
  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems
  • Profile and tune low level CPU and memory performance
Requirements
  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning
  • Deep understanding of computational complexity and memory allocation
  • Track record of designing systems that scale under real production load
  • Experience building and using observability to drive performance and reliability decisions
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
  • Experience with AI/ML inference or training infrastructure
  • Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)
  • Background in building multi-tenant compute platforms
  • Understanding of networking fundamentals and performance characteristics
  • Familiarity with GPU workload characteristics and scheduling constraints
Compensation
  • $180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff)
Location
  • San Francisco, CA (willing to consider remote for Senior and Staff levels)

What we offer at fal
  • Interesting and challenging work

  • A lot of learning and growth opportunities

  • We are currently hiring in downtown San Francisco.

  • We offer relocation assistance to San Francisco.

  • Health, dental, and vision insurance (US)

  • Regular team events and offsites

Similar Jobs

8 Days Ago
Hybrid
Mountain View, CA, USA
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Build and optimize distributed compute infrastructure for large-scale analytical and AI workloads. Responsibilities include improving query execution, scheduling, resource allocation, reliability, workload routing, and compute efficiency across Spark and related systems. The role involves debugging performance bottlenecks, optimizing joins, scans, shuffles, caching, partitioning, and memory usage, and working with lakehouse formats and cloud object storage. Candidates will implement workload optimization algorithms and may contribute to open source or research.
Top Skills: Adaptive Query ExecutionAmazon EmrAmazon S3Apache FlinkApache HiveApache HudiApache IcebergSparkAws GlueAzure Data Lake StorageC++CatalystDatabricksDatafusionDelta LakeDuckdbGoGoogle Cloud StorageJavaOrcParquetPrestoRustScalaSnowflakeSpark SqlTrinoVelox
16 Days Ago
In-Office
Palo Alto, CA, USA
158K-237K Annually
Mid level
158K-237K Annually
Mid level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Design, develop, and deliver distributed systems solutions for Rubrik’s Atlas data platform. Responsibilities include defining architectural principles, prototyping, coding production-quality systems, and guiding multiple teams. Work spans cloud-backed file systems, optimized data formats, deduplication, compression, resiliency, scalability, performance tooling, asynchronous programming, snapshotting, data integrity, security, and immutability.
Top Skills: AWSAzureAzure BlobC++Distributed File SystemsDistributed SystemsEbpfEbsEc2Ext4GCPHddKubernetesNfsOracleRdsS3SmbSsd
21 Days Ago
Hybrid
Santa Clara, CA, USA
150K-262K Annually
Senior level
150K-262K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design, build, and operate cloud-native data platform components and services. Lead and deliver well-scoped projects, implement operators/controllers and automation, ensure reliability and observability, participate in on-call, and mentor junior engineers.
Top Skills: AIAWSAzureCi/CdCniContainersDashboardsGCPGitGitopsGoInfrastructure-As-CodeKubernetesKubernetes ControllersMetricsMtlsObservabilityOperatorsSecrets ManagementService MeshTracingWorkload Identity

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account