Inworld AI Logo

Inworld AI

Staff / Principal Machine Learning Engineer, Serving - USA

Reposted One Month Ago
Be an Early Applicant
Hybrid
Mountain View, CA, USA
270K-500K Annually
Expert/Leader
Hybrid
Mountain View, CA, USA
270K-500K Annually
Expert/Leader
This role involves optimizing inference performance, managing high-performance systems, scaling models, and full-cycle ownership from research to production in a fast-paced environment.
The summary above was generated by AI

About Inworld

Inworld is a research lab and inference provider focused on realtime AI for consumer-facing applications. We build first-party speech models, serve LLMs, and run the inference behind modular APIs designed for high-volume, realtime workloads.

Hundreds of millions of users interact with Inworld powered apps every day and we serve over 10 trillion LLM tokens per month. Our models and infrastructure support consumer applications across companions, healthcare, fitness, education, media, and more. Our work spans model research, realtime inference, large-scale serving infrastructure, and the APIs developers use to bring these capabilities into production.

We’ve raised more than $125M from Lightspeed Venture Partners, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta, Stanford, and others. Our technology has powered experiences from companies including NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella, and Bible Chat. Inworld has also been recognized by CB Insights as one of the 100 most promising AI companies globally and named one of LinkedIn’s Top 10 Startups in the USA.

Who We're Looking For

A year ago, reliably working agentic systems and sub-second multimodal inference at scale barely existed. Nobody has a decade of experience here. So we're not screening for a resume template — we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood.

Experience We Find Useful

You don't need all of this. But you need enough to make a case.

  • Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.

  • Model Acceleration. Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding.

  • High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. You know how to profile code and squeeze every ounce of performance out of NVIDIA GPUs.

  • Distributed Systems & Scaling. Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections.

  • Public work. Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups.

  • Full-cycle ownership. You can take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production.

  • Background. PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.

Who Thrives Here

  • You don’t need a roadmap to start walking; you’re comfortable picking a direction and building the map as you go.

  • You believe engineering isn't finished until it’s shipped and stable. You have a bias for impact over purely theoretical optimizations.

  • You don't just ship code; you obsess over the why. You’re the first to question an architecture if you think there’s a better way to solve the core latency or throughput problem.

  • You aren't satisfied with "the PM said so." You thrive on deep context and want to understand the fundamental logic behind every decision we make.

What Working Here Is Like

We hand you unclear problems and expect you to make them clear. We value engineers who say "I don't know yet" and then design the benchmark or prototype that finds out. We treat performance, latency, and reliability as first-class product features, not a box to check before launch. Impact comes before everything else, though we support sharing work and open-source contributions that move the field forward. Your work should be visible. Flat structure, fast iterations, minimal process theater.

We believe in the power of in-person collaboration to solve the hardest problems and foster a strong team culture. We offer relocation assistance and look forward to you joining us in our Mountain View office.

The base salary range for this full-time position is $270,000 - $500,000+ bonus + equity + benefits.

Inworld Jobs Privacy

HQ

Inworld AI Mountain View, California, USA Office

1975 W El Camino Real, Mountain View, CA, United States, 94040

Similar Jobs

4 Minutes Ago
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Financial Services
Provides executive administrative support through complex calendar management, call screening, meeting and event coordination, domestic and international travel arrangements, invoice and expense processing, onboarding and offboarding support, document maintenance, and preparation of professional communications, spreadsheets, and presentations. The role requires discretion, strong organization, communication skills, Microsoft Office proficiency, and five days of on-site work.
Top Skills: MS Office
8 Minutes Ago
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • HR Tech • Information Technology • Machine Learning • Software • App development • Industrial
Own the full order-to-cash function, including invoicing and payment policy, credit decisions, collections, AR aging, customer holds, and outsourced collections management. Partner with Sales, Finance, Product, Engineering, and Operations to resolve aged receivables and automate credit, invoicing, PO matching, customer-portal delivery, cash application, and VMS billing workflows. Lead roadmap requirements, rollout, KPIs, and process improvements in a high-growth environment.
Top Skills: Ai-Native Finance Automation ToolsAr DashboardsCustomer PortalsNetSuiteQuickbooks Online (Qbo)Vms/Msp Billing Systems
Junior
Artificial Intelligence • Healthtech • Logistics • Social Impact • Software • Telehealth
Engages patients through high-volume inbound and outbound calls, educates them about ordered at-home healthcare services, schedules appointments, answers questions, and escalates concerns. The role requires approximately 150–200 outbound calls and 80–100 patient conversations daily, collaboration with healthcare professionals, and proficiency with EHR and healthcare software. Zendesk and Five9 experience are preferred.
Top Skills: Auto-Dialer SystemsElectronic Health Records (Ehr)Five9Zendesk

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account