Together AI Logo

Together AI

Senior Machine Learning Engineer, Voice AI

Reposted One Month Ago
In-Office
San Francisco, CA, USA
160K-230K Annually
Senior level
In-Office
San Francisco, CA, USA
160K-230K Annually
Senior level
This role involves optimizing model serving layers for voice AI applications, working with inference engines, and improving performance for STT and TTS systems.
The summary above was generated by AI
About the Role

Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform powers production-grade, real-time voice agents and applications — serving speech-to-text and text-to-speech models with best-in-class latency and reliability.

We're looking for a Senior ML Engineer to drive the model serving layer for voice workloads. You'll work hands-on with inference engines like TRT-LLM and SGLang to optimize how we serve models like Whisper, Parakeet, Orpheus, and Kokoro — pushing latency and throughput to the frontier. You'll profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly.

This is a foundational hire on a small, high-impact team. Voice inference has unique challenges — streaming audio, tokenization, real-time latency budgets — that require dedicated ML engineering focus. You'll shape how Together serves voice models as the industry moves from pipeline architectures (ASR → LLM → TTS) toward end-to-end speech-to-speech.

  • Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech.
  • Work directly with state-of-the-art accelerators (H100s, H200s, B200s) to optimize voice model inference.
  • Collaborate with model partners (Cartesia, Deepgram, Rime, and others) to bring their models to production on Together's infrastructure.
  • Build quality evaluation frameworks that guide model selection for customers and inform the roadmap.
  • Join a small, early-stage team with outsized impact on a fast-growing product area.
Responsibilities
  • Optimize inference performance for voice models (STT, TTS, speech-to-speech) — targeting best-in-class TTFB, throughput, and GPU utilization across our curated model set.
  • Productionize voice models on serverless and dedicated endpoints, including batching strategies, streaming inference, and memory management tailored to audio workloads.
  • Build and maintain a voice model evaluation framework — measuring WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation accuracy for TTS.
  • Enable new model architectures in our serving stack as the field evolves, including audio-native LLMs, codec-based models (SNAC), and speech-to-speech systems.
  • Collaborate with model partners to integrate and optimize their models (Cartesia, Deepgram, Rime, and others) running on Together's infrastructure.
  • Profile and debug performance across the full inference stack — from GPU kernels to framework-level bottlenecks — and ship measurable improvements.
  • Work with the platform engineering side of the team to ensure the serving layer meets the latency and reliability requirements of real-time voice APIs.
  • Contribute to voice model fine-tuning capabilities (STT and TTS) as we enable customers to build differentiated voice experiences on Together.
  • Lay the groundwork for multiple new products down the line.
Requirements
  • 5+ years of experience in ML engineering, with a focus on model serving, inference optimization, or ML infrastructure.
  • Hands-on experience with LLM serving engines (vLLM, SGLang, TensorRT-LLM, or similar) — comfortable reading and modifying engine internals, not just using APIs.
  • Strong proficiency in Python and PyTorch; experience with GPU profiling and optimization (CUDA, memory management, kernel-level debugging).
  • Track record of shipping ML systems to production with measurable performance improvements.
  • Strong product sense — you think about what developers building voice apps actually need, not just what's technically interesting.
  • Comfort working on a small, early-stage team where you'll wear multiple hats and move fast.
  • Experience with speech and audio ML (ASR, TTS architectures, audio signal processing) is a strong plus but not required — you can learn this quickly if you have strong ML engineering fundamentals.
  • Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC) is a plus.
  • Experience training or fine-tuning speech models is a plus.
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience
About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200,000 - $260,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Together AI San Francisco, California, USA Office

584 Castro St, #2050, San Francisco, California , United States, 94114

Similar Jobs

8 Minutes Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
156K-234K Annually
Senior level
156K-234K Annually
Senior level
Artificial Intelligence • Cloud • Software
Own full-cycle recruiting for go-to-market functions, including sourcing, pipeline development, candidate screening, interview management, offer negotiation, and closing. Partner with hiring managers to define requirements, calibrate candidates, and improve hiring processes. Track recruiting metrics, report progress and risks, and use AI and recruiting technology to improve efficiency and candidate experience.
Top Skills: Ai Recruiting ToolsRecruiting Technology
59 Minutes Ago
Hybrid
Santa Clara, CA, USA
127K-222K Annually
Senior level
127K-222K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads global employee communications, cultural moments, engagement campaigns, and company events. Develops narratives around business priorities and AI, partners across communications and business teams, manages programs from strategy through execution, supports live events, and uses engagement data to improve employee experiences.
Top Skills: AI
An Hour Ago
Remote or Hybrid
United States
Senior level
Senior level
Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Leads configuration data management and release activities for aerospace programs. Reviews engineering drawings, EBOMs, and ECOs; maintains configuration plans; performs audits and root-cause analysis; supports PLM tools, upgrades, and testing; ensures compliance with aerospace standards and contractual requirements; provides reporting and stakeholder support; and mentors team members while driving process improvements.
Top Skills: ClearcaseClearquestProduct Data Management (Pdm)Product Lifecycle Management (Plm)SalesforceTeamcenter

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account