FriendliAI Logo

FriendliAI

Software Engineer – Cloud Infrastructure

Posted 8 Days Ago
Hybrid
San Francisco, CA, USA
Senior level
Hybrid
San Francisco, CA, USA
Senior level
Design, build, and operate a multi-cluster, multi-tenant Kubernetes fleet for GPU inference. Extend Kubernetes with controllers/CRDs, implement GPU scheduling and autoscaling, own the network data plane, design cross-cluster connectivity and service mesh, drive reliability/SLOs, deliver IaC with Terraform/Helm/GitOps, and collaborate with platform, SRE, and security teams.
The summary above was generated by AI
About the job

FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.

Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.

Key Responsibilities

Cluster Architecture

  • Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades.

  • Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks.

  • Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants.

  • Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.

Networking

  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.

  • Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.

  • Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes.

Reliability & Collaboration

  • Define SLOs for platform-critical systems and lead post-incident hardening.

  • Deliver infrastructure as code with Terraform, Helm, and GitOps.

  • Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.

Qualifications
  • 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production.

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.

  • Proven experience operating large-scale, high-traffic network services in production.

  • Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd.

  • Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.

  • Proficiency with AWS, Terraform, Helm, and Ansible.

  • Programming skills in Go or Python, with the ability to build infrastructure tooling and automation.

  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.

  • Clear written and verbal communication, including the ability to document architectural decisions for other engineers.

Preferred Experience
  • Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.

  • Cilium and eBPF, including kube-proxy replacement or upstream contributions.

  • Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.

  • GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA).

  • High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning.

  • Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations.

  • Contributions to Kubernetes, Cilium, Istio, or other CNCF projects.


Benefits
  • Flexible working hours

  • Daily lunch and dinner provided; unlimited snacks and beverages

  • Supportive and highly collaborative work environment

  • Health check-up support and top-tier equipment/hardware support

  • A front-row seat to the generative AI infrastructure revolution

  • Competitive compensation, startup equity, health insurance, and other benefits.

About FriendliAI

FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling.

We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

Similar Jobs

6 Days Ago
In-Office
Santa Clara, CA, USA
124K-242K Annually
Junior
124K-242K Annually
Junior
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop AI agents and software fixes to automate defect intake, triage, assignment, and pipeline handling. Collaborate with QA to reproduce defects, implement timely fixes, and improve AI-driven feedback loops for open-source cloud development.
Top Skills: AIAi AgentsDgx CloudGitLlms
15 Days Ago
In-Office or Remote
2 Locations
Senior level
Senior level
Aerospace • Automation
Design, build, and maintain cloud-based data pipelines and infrastructure on AWS, integrating on-prem and cloud flows. Implement IaC, CI/CD, containerization, observability, and automation to ensure high-throughput, reliable data ingestion and processing for autonomous systems.
Top Skills: AWSCircleCICloudFormationDockerGitGithub ActionsGrpcJenkinsKafkaKubernetesPythonTerraform
15 Days Ago
In-Office
San Francisco, CA, USA
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Lead design, development, and operation of cloud infrastructure and SRE-focused systems. Own medium-to-large infrastructure projects, build resilient platforms, and drive cross-team technical delivery. Mentor engineers, define SLOs, reduce failure domains, and build tooling for automated, secure CI/CD and production reliability.
Top Skills: ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account