FriendliAI

United States
34 Total Employees
Year Founded: 2021

Jobs at FriendliAI

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

21 Days AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Own the vision, strategy, and roadmap for FriendliAI’s AI inference platform, including model serving, deployment, orchestration, APIs, and developer capabilities. Partner with engineering, research, enterprise customers, and go-to-market teams to improve inference performance, scalability, reliability, usability, and cost. Lead product initiatives from discovery through launch, establish product planning and measurement practices, and mentor Product Managers and Designers.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and operate a multi-cluster, multi-tenant Kubernetes fleet for GPU inference. Extend Kubernetes with controllers/CRDs, implement GPU scheduling and autoscaling, own the network data plane, design cross-cluster connectivity and service mesh, drive reliability/SLOs, deliver IaC with Terraform/Helm/GitOps, and collaborate with platform, SRE, and security teams.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Own and evolve core backend microservices for an AI inference platform: build production-grade APIs, multi-tenant SaaS features (auth, RBAC, billing), design OLTP/OLAP data models, collaborate on multi-cloud orchestration, ensure reliability/performance, and drive engineering quality through testing and CI/CD.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and maintain agent APIs and production agent applications for document understanding, advanced RAG, and customer support automation. Integrate open-source models, collaborate with backend and infra for deployment and monitoring, and ensure APIs are robust, scalable, and developer-friendly.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design and optimize high-performance GPU kernels (GEMM, attention, routing) for AI inference across NVIDIA and AMD GPUs. Implement CUDA/C++ and low-level assembly code, build reduced-precision/quantized (FP8/FP4) kernels, benchmark cross-vendor performance, contribute to internal GPU libraries, accelerate multi-modal pipelines, and integrate next-generation GPU features into production.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and maintain a scalable web platform and APIs for deploying and monitoring multimodal AI models and agent workflows. Collaborate with product, infrastructure, and design teams to optimize performance, ensure reliability, drive CI/CD and testing, and contribute to long-term architecture decisions for a cloud-native, multi-tenant SaaS system.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Lead end-to-end enterprise sales for FriendliAI's AI inference platform: generate pipeline, close high-value deals, run technical POCs, engage AI/ML communities, collaborate with engineering, and inform product roadmap.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, implement, and optimize GPU kernels, kernel compiler, memory planner, and runtime for low-latency generative AI inference. Analyze performance bottlenecks across hardware and software, collaborate with infrastructure teams, and maintain production profiling, benchmarking, and validation tooling while supporting new model architectures and multi-GPU strategies.
One Month AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, deploy, and operate large-scale LLM and multimodal inference architectures. Work hands-on with customer engineering teams to containerize, scale, monitor, and troubleshoot GPU-based inference workloads across Kubernetes, CI/CD, and hybrid/on-prem environments. Create Helm charts, Terraform modules, and observability tooling while delivering workshops and platform reliability insights.