fal Logo

fal

Software Engineer, Infrastructure

Reposted 14 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
180K-250K Annually
Senior level
In-Office
San Francisco, CA, USA
180K-250K Annually
Senior level
Build and maintain Python-based fleet management and server tooling for thousands of GPU servers, automate provisioning, health monitoring, diagnostics, and recovery, create metrics/dashboards, enforce OS security, manage storage, tune Linux for AI workloads, and drive resolutions with partners.
The summary above was generated by AI

You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including  provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.

Key responsibilities
  • Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc
  • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting
  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)
  • Leverage AI to an extreme level to build tools and automate alerting and recovery
  • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation
  • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage
  • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)
  • Develop a suite of automated error detection and recovery processes
  • Work with partners to solve technical issues
Requirements
  • 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)
  • Strong software engineering skills in Python; you write production tooling, not scripts
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
  • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
  • Experience building internal tools or dashboards for infrastructure visibility
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
  • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)
  • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2
  • Experience with AMD GPUs
  • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)
  • Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001)
Compensation
  • $180,000-250,000 plus equity + benefits
Location
  • San Francisco, CA (we are open to remote in the US for Senior and Staff levels)

What we offer at fal
  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We are offering relocation assistance to San Francisco.
  • We offer relocation assistance to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

Similar Jobs

4 Days Ago
Easy Apply
Hybrid
San Jose, CA, USA
Easy Apply
158K-225K Annually
Senior level
158K-225K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Architect, develop, and maintain scalable system test automation infrastructure for ZIA Core Datapath. Build automation suites validating cloud firewall, tunneling, load balancing, IPS/IDS, routing, and switching technologies. Perform functional, integration, performance, and scale testing; troubleshoot complex network, HTTP/SSL, and firewall issues; optimize regression cycles; and mentor engineers.
Top Skills: Ai-Assisted CodingDhcpDnsDynamic AnalysisFuzzingGreHttp(S)Ips/IdsIpsecL3/L4/L7 FilteringLoad BalancersLoad TestingPerformance TestingPkiPythonRoutersRouting And SwitchingStatic AnalysisTcp/IpTlsTraffic GeneratorsVpn
5 Days Ago
Hybrid
Palo Alto, CA, USA
133K-235K Annually
Junior
133K-235K Annually
Junior
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Build and optimize large-scale machine learning infrastructure for content retrieval and recommendation. Responsibilities include developing feature generation and serving pipelines, high-performance inference systems, cloud-based training and evaluation infrastructure, and data management systems. The role partners with ML engineers to deploy models, improve reliability and efficiency, and operate highly available distributed systems using Java, Go, C++, and Python.
Top Skills: C++Caffe2FlinkGoJavaPythonPyTorchRayScikit-LearnSparkSpark MlTensorFlow
10 Days Ago
Hybrid
Mountain View, CA, USA
143K-243K Annually
Senior level
143K-243K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design, build, and evolve reliable, performant, and secure core infrastructure platform components. Own multiple platform areas, influence the roadmap, develop cross-cutting solutions with engineering teams, deliver interdependent work on schedule, and drive continuous platform improvement. Participate in code and design reviews while collaborating with product and cross-functional stakeholders. The role requires building API-based platforms and maintaining microservices using Python, Golang, Java, or C++.
Top Skills: C++GoJavaMicroservicesPythonRestRpc

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account