F5 Logo

F5

SRE, AI Inference Engineer

Posted 20 Days Ago
Be an Early Applicant
In-Office
San Jose, CA, USA
177K-265K Annually
Entry level
In-Office
San Jose, CA, USA
177K-265K Annually
Entry level
Build and optimize high-performance LLM inference systems across data centers and edge devices. Develop inference engines, accelerate models on GPUs and AI hardware, design Kubernetes-based autoscaling architectures, and implement observability, load testing, and reliability frameworks. Improve latency, throughput, cost efficiency, and scalability across real-time and batch AI workloads while collaborating with hardware and cross-functional engineering teams.
The summary above was generated by AI

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. 
 

Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.

Job Description

The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices—with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy.  

This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance. 

Key Responsibilities 

High-Performance AI Serving 

  • Build and maintain robust inference engines using tools like vLLMTGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale.  

  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications. 

Hardware Acceleration and Optimization 

  • Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs.  

  • Collaborate with hardware teams to maximize utilization and performance across various computational environments. 

Inference Orchestration and Scalability 

  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration.  

  • Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability. 

Performance Monitoring and Observability 

  • Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs).  

  • Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale. 

Technical Requirements 

Required Skills: 

  • Programming Languages: Proficiency in programming languages such as PythonC++Rust, or Golang specifically for high-performance AI workflows.  

  • Inference Tools: Proven hands-on experience with tools like vLLMTensorRTLlama.cpp, and Ollama for inference development and optimization.  

  • Infrastructure Expertise: Strong familiarity with infrastructure technologies, including DockerKubernetes, and cloud platforms such as AWSGCP, and Azure.  

  • Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs. 

Preferred Experience: 

  • Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention.  

  • Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels).  

  • Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges.  

  • Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts. 

Success Metrics (KPIs): 

  • Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) and maximum tokens per second.  

  • Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resource utilization and hardware optimization.  

  • Scalability: Maintain system stability and reliability during traffic spikes, ensuring performance consistency across environments.  

  • Throughput Maximization: Deploy models optimized for peak hardware usage and maximized process throughput. 

Why Join F5? 

F5 empowers you to push boundaries in AI optimization and high-performance engineering. Joining our team means:  

  • Collaborating with cutting-edge technologies and hardware solutions to support real-time AI applications.  

  • Advancing your career in a fast-paced, multidisciplinary environment focused on innovation, scalability, and problem-solving.  

  • Driving transformative projects that deliver real-time AI reliability to global customers while maintaining cost and efficiency standards.  

  • Working on advanced MLOps solutions that seamlessly scale enterprise AI systems and shape the future of intelligent deployment. 

What Success Looks Like: 

As an AI Inference Engineer at F5, success is measured by your ability to:  

  • Combine technical expertise and problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems.  

  • Collaborate efficiently across cross-functional teams, participating in knowledge sharing and system refinement.  

  • Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals.  

  • Translate complex AI and inference workflows into practical solutions that align with F5's strategic objectives. 

The base pay range per annum for this position is: $176,600 - $265,000

F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5’s differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change. You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5’s benefits can be found at the following link: https://www.f5.com/company/careers/benefits. F5 reserves the right to change or terminate any benefit plan without notice.

#LI-ZB1

The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.

Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or @myworkday.com).

Equal Employment Opportunity

It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination.  F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting [email protected].

F5 San Jose, California, USA Office

90 Rio Robles, San Jose, CA, United States, 95134

Similar Jobs

26 Minutes Ago
Remote or Hybrid
United States
98K-164K Annually
Expert/Leader
98K-164K Annually
Expert/Leader
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Leads a team of Forward Deployed Engineers while supporting sales, customer discovery, solution architecture, foundational implementations, and proof-of-value engagements. The role translates customer requirements into SailPoint solutions, advises on product architecture and best practices, collaborates across sales, product, services, and customer success, and develops AI-powered agentic solutions for demonstrations and customer outcomes. It requires significant enterprise software implementation or presales experience, technical leadership, SaaS architecture knowledge, and client-facing consulting expertise.
Top Skills: Active DirectoryAgentic AiAngularArtificial IntelligenceAWSAzureCassandraClaude CodeCursorGCPGithub CopilotGoogle AntigravityJ2EeJavaJavaScriptJSONKiroLdapLinuxMicrosoft Sql ServerMongoDBMySQLNode.jsOne IdentityOpenai CodexOracleOracle Identity ManagerPeoplesoftPowershellPythonReactRedisRsa AveksaSailpointSAPSaviyntServicenowSQLSybaseTypescriptUnixWindowsXML
31 Minutes Ago
Remote or Hybrid
USA
20-23 Hourly
Entry level
20-23 Hourly
Entry level
Fintech • Healthtech • HR Tech • Information Technology • Financial Services • Telehealth
Provides empathetic phone and chat support to families navigating loss, disability leave, and other difficult life events. Responsibilities include creating care plans, researching resources and service providers, answering logistical questions, documenting interactions, meeting service levels, escalating sensitive issues, maintaining knowledge resources, and sharing user insights. The role requires independent work, strong communication, organization, problem-solving, adaptability, and comfort with technology and sensitive data.
Top Skills: Google SuiteSlackZendesk
37 Minutes Ago
In-Office
197K-202K Annually
Senior level
197K-202K Annually
Senior level
Other • Utilities
Gather and engineer complex datasets from internal and external sources; develop ETL pipelines, dashboards, visualizations, forecasts, and advanced analyses using SQL, Python, Power BI, Tableau, Snowflake, and Databricks. Lead requirements gathering, prioritize analytics projects, define KPIs and metrics, create roadmaps, ensure data quality, and communicate strategic insights and recommendations to senior leadership.
Top Skills: DatabricksDaxExcelPower BIPythonSnowflakeSQLTableau

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account