Together AI Jobs

Senior Machine Learning Engineer, Voice AI

Together AI

Senior Machine Learning Engineer, Voice AI

Reposted 8 Days Ago

In-Office

San Francisco, CA, USA

160K-230K Annually

Senior level

In-Office

San Francisco, CA, USA

160K-230K Annually

Senior level

This role involves optimizing model serving layers for voice AI applications, working with inference engines, and improving performance for STT and TTS systems.

The summary above was generated by AI

About the Role

Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform powers production-grade, real-time voice agents and applications — serving speech-to-text and text-to-speech models with best-in-class latency and reliability.

We're looking for a Senior ML Engineer to drive the model serving layer for voice workloads. You'll work hands-on with inference engines like TRT-LLM and SGLang to optimize how we serve models like Whisper, Parakeet, Orpheus, and Kokoro — pushing latency and throughput to the frontier. You'll profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly.

This is a foundational hire on a small, high-impact team. Voice inference has unique challenges — streaming audio, tokenization, real-time latency budgets — that require dedicated ML engineering focus. You'll shape how Together serves voice models as the industry moves from pipeline architectures (ASR → LLM → TTS) toward end-to-end speech-to-speech.

Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech.
Work directly with state-of-the-art accelerators (H100s, H200s, B200s) to optimize voice model inference.
Collaborate with model partners (Cartesia, Deepgram, Rime, and others) to bring their models to production on Together's infrastructure.
Build quality evaluation frameworks that guide model selection for customers and inform the roadmap.
Join a small, early-stage team with outsized impact on a fast-growing product area.

Responsibilities

Optimize inference performance for voice models (STT, TTS, speech-to-speech) — targeting best-in-class TTFB, throughput, and GPU utilization across our curated model set.
Productionize voice models on serverless and dedicated endpoints, including batching strategies, streaming inference, and memory management tailored to audio workloads.
Build and maintain a voice model evaluation framework — measuring WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation accuracy for TTS.
Enable new model architectures in our serving stack as the field evolves, including audio-native LLMs, codec-based models (SNAC), and speech-to-speech systems.
Collaborate with model partners to integrate and optimize their models (Cartesia, Deepgram, Rime, and others) running on Together's infrastructure.
Profile and debug performance across the full inference stack — from GPU kernels to framework-level bottlenecks — and ship measurable improvements.
Work with the platform engineering side of the team to ensure the serving layer meets the latency and reliability requirements of real-time voice APIs.
Contribute to voice model fine-tuning capabilities (STT and TTS) as we enable customers to build differentiated voice experiences on Together.
Lay the groundwork for multiple new products down the line.

Requirements

5+ years of experience in ML engineering, with a focus on model serving, inference optimization, or ML infrastructure.
Hands-on experience with LLM serving engines (vLLM, SGLang, TensorRT-LLM, or similar) — comfortable reading and modifying engine internals, not just using APIs.
Strong proficiency in Python and PyTorch; experience with GPU profiling and optimization (CUDA, memory management, kernel-level debugging).
Track record of shipping ML systems to production with measurable performance improvements.
Strong product sense — you think about what developers building voice apps actually need, not just what's technically interesting.
Comfort working on a small, early-stage team where you'll wear multiple hats and move fast.
Experience with speech and audio ML (ASR, TTS architectures, audio signal processing) is a strong plus but not required — you can learn this quickly if you have strong ML engineering fundamentals.
Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC) is a plus.
Experience training or fine-tuning speech models is a plus.
Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200,000 - $260,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy

584 Castro St, #2050, San Francisco, California , United States, 94114

Similar Jobs

GEICO

Staff Engineer

5 Days Ago

In-Office

Palo Alto, CA, USA

130K-260K Annually

Expert/Leader

130K-260K Annually

Expert/Leader

Insurance

As a Senior Staff Engineer, you will lead engineering teams, provide technical solutions, mentor junior members, and drive quality in enterprise applications with focus on voice technologies.

Top Skills: .NetAWSAzure DevopsC#JavaNode.jsPowershellPython

Square

Account Executive

28 Minutes Ago

Hybrid

123K-223K Annually

Mid level

123K-223K Annually

Mid level

eCommerce • Fintech • Hardware • Payments • Software • Financial Services

Field-driven full-cycle sales role responsible for building pipeline, running in-person demos, and closing deals across Squares product suite. Spend ~80% of time in-market with 50-60 targeted visits weekly, manage pipeline in Salesforce, establish local partnerships, and exceed quota while onboarding and supporting new sellers.

Top Skills: AfterpayCash AppLoyalty PlatformsPayment Processing SystemsPayroll SystemsSalesforceSquareTime Management Software

Wipfli

Data Architect

59 Minutes Ago

Remote or Hybrid

United States

142K-191K Annually

Senior level

142K-191K Annually

Senior level

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services

The Enterprise Data Architect will design and govern data architecture, create enterprise data models, and ensure data consistency across various systems.

Top Skills: DatabricksDynamics 365Microsoft FabricPower BIWorkday

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Together AI

Senior Machine Learning Engineer, Voice AI

Together AI San Francisco, California, USA Office

Similar Jobs

Staff Engineer

Account Executive

Data Architect

What you need to know about the San Francisco Tech Scene

Key Facts About San Francisco Tech