FriendliAI Logo

FriendliAI

Solutions Architect - AI Inference Specialist

Reposted One Month Ago
Be an Early Applicant
Hybrid
San Francisco, CA, USA
Mid level
Hybrid
San Francisco, CA, USA
Mid level
Design, deploy, and operate large-scale LLM and multimodal inference architectures. Work hands-on with customer engineering teams to containerize, scale, monitor, and troubleshoot GPU-based inference workloads across Kubernetes, CI/CD, and hybrid/on-prem environments. Create Helm charts, Terraform modules, and observability tooling while delivering workshops and platform reliability insights.
The summary above was generated by AI

About the job

FriendliAI is seeking a Solution Architect to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container.

Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product.

You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position.

Key Responsibilities

  • Design and implement large-scale deployment architectures for LLM and multimodal inference

  • Deploy and manage containerized workloads across Kubernetes clusters

  • Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed

  • Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows

  • Develop scripts, Helm charts, and Terraform modules that simplify repeated deployments

  • Contribute field insights to shape our platform reliability, observability, and scaling strategies

  • Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices.

Qualifications

  • 3+ years of experience in cloud infrastructure, DevOps, or reliability engineering

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

  • Proficiency with Kubernetes, Docker, Terraform, and Helm

  • Strong foundation in distributed systems, networking, and performance tuning

  • Experience with GPU-based computing and generative AI model serving workloads

  • Strong technical background in backend systems or AI tooling

  • Experience operating workloads on AWS, GCP, or OCI

  • Excellent problem-solving and debugging skills in real-world environments

Preferred Experience

  • Experience deploying large models (LLMs, diffusion models) on GPUs or clusters

  • Familiarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference)

  • Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL)

  • Understanding of networking security and compliance frameworks (e.g., SOC 2)

  • Experience supporting on-prem or hybrid-cloud deployments

Benefits

  • A front-row seat to the generative AI infrastructure revolution

  • Competitive compensation and benefits package

  • Daily lunch and dinner provided; unlimited snacks and beverages

  • Health check-up and top-tier hardware support

  • Flexible working hours and a highly collaborative environment

About us

FriendliAI is building the next-generation AI inference platform that accelerates the deployment of large language and multimodal models with unmatched performance and efficiency. Our infrastructure powers high-throughput, low-latency workloads for global organizations and integrates directly with Hugging Face, providing instant access to over 600,000 open-source models. We are on a mission to deliver the world’s best platform for AI inference.

Similar Jobs

An Hour Ago
In-Office
106K-232K Annually
Senior level
106K-232K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Designs, integrates, tests, debugs, and leads development of embedded flight software for satellite and payload systems. Responsibilities include software-hardware integration, requirements analysis, cyber vulnerability analysis, cyber monitoring algorithms, configuration automation, documentation, quality assurance, and delivery coordination. The role interfaces with multidisciplinary engineering teams and supports safety, security, and performance objectives for commercial and government space programs.
Top Skills: BitbucketCC++ConfluenceCyber Monitoring AlgorithmsDevOpsEmbedded SystemsGitlabJavaJIRAPythonReal-Time SoftwareWireshark
An Hour Ago
In-Office
177K-239K Annually
Expert/Leader
177K-239K Annually
Expert/Leader
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Leads electrophysics and radar engineering for air and space payloads across the product lifecycle. Responsibilities include radar performance and trade studies, signal and image processing, RWR integration, mission design, payload effectiveness analysis, technical leadership of teams and subcontractors, anomaly resolution, and capability roadmap development. The role requires advanced radar or signal-intelligence expertise, systems integration experience, and active Top Secret/SCI clearance.
Top Skills: Image ProcessingMatlabRadar SystemsRadar Warning Receivers (Rwr)Rf SystemsRfi MitigationSignal ProcessingStkSynthetic Aperture Radar (Sar)
An Hour Ago
In-Office
120K-198K Annually
Senior level
120K-198K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Design and validate ASIC and mixed-signal subsystem architectures for space-based digital communication systems. Responsibilities include requirements derivation, verification planning, ADC/DAC performance analysis, technical trade studies, supplier coordination, hardware and software integration, qualification, unit sell-off, and risk mitigation. The role supports the full subsystem lifecycle and requires collaboration across engineering, program management, suppliers, and customers.
Top Skills: AdcAsicCDacI2CJesdJtagMatlabPythonSerdesSpiVlsi

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account