xAI Logo

xAI

Software Engineer - Training/Inference (C++)

Reposted 7 Days Ago
Be an Early Applicant
In-Office
Palo Alto, CA, USA
180K-440K Annually
Senior level
In-Office
Palo Alto, CA, USA
180K-440K Annually
Senior level
Design and optimize large-scale, low-latency model serving systems end-to-end. Implement distributed infrastructure (batching, caching, load balancing, autoscaling), accelerate GPU inference (kernels, quantization, codegen), build reliable high-concurrency serving, and create tooling and CI/CD for deployment and benchmarking.
The summary above was generated by AI

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

About the Role:
  • We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability.
  • As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency).
  • This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale

Responsibilities: 

  • Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
  • Optimize latency and throughput of model inference under real production workloads.
  • Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency.
  • Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation).
  • Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels.
  • Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
  • Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.
BASIC QUALIFICATIONS:
  • Deep low-level systems programming (C/C++ or Rust)
  • Experience with large-scale, high-concurrent production serving.
  • Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
  • Strong background in system optimizations: batching, caching, load balancing, parallelism.
  • Low-level inference optimizations: GPU kernels, code generation.
  • Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.
  • Experience with testing, benchmarking, and reliability of inference services.
  • Experience designing and implementing CI/CD infrastructure for inference.
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

HQ

xAI Palo Alto, California, USA Office

1450 Page Mill Road, Palo Alto, CA, United States

xAI San Francisco, California, USA Office

3180 18th St., San Francisco, CA, United States

Similar Jobs

An Hour Ago
Easy Apply
Hybrid
Easy Apply
Mid level
Mid level
Fintech • Financial Services
Lead and support day-to-day post-origination loan servicing across real estate, commercial, and consumer loans. Serve as first-level escalation for teammates and members, perform quality control reviews, ensure compliance with lending regulations, assist with audits, coach and develop staff, manage servicing workflows (payoffs, lien releases, escrow, insurance tracking), and partner cross-functionally to resolve servicing issues and improve processes.
Top Skills: EncompassMeridian LinkMS OfficeSymitar
2 Hours Ago
In-Office
San Jose, CA, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Hardware • Information Technology • Machine Learning
The Staff Cloud Solutions Architect will lead technology strategy for cloud and enterprise accounts, build customer relationships, guide solution design, and influence product roadmaps.
Top Skills: Ai InfrastructureData Center EnvironmentsDramMemory TechnologiesServer ArchitecturesSsdStorage Technologies
2 Hours Ago
In-Office
San Jose, CA, USA
142K-242K Annually
Senior level
142K-242K Annually
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Act as Micron's technical advisor for datacenter customers to drive DRAM/HBM/SSD design wins, resolve system configuration and performance issues, guide memory architecture and optimization, support validation/lab testing and deployments, and collaborate across sales, product, and engineering teams.
Top Skills: Ai ToolsBmcDramEdacEmcHbmLogic AnalyzersMemory InterfacesOscilloscopesServer ArchitectureSsd

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account