Together AI Logo

Together AI

Machine Learning Engineer - Inference

Reposted One Month Ago
In-Office
San Francisco, CA, USA
160K-230K Annually
Mid level
In-Office
San Francisco, CA, USA
160K-230K Annually
Mid level
Design and build production systems for AI inference, optimize runtime services, and collaborate on AI solutions with researchers and engineers.
The summary above was generated by AI
About the Role

Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and ensuring they run efficiently and effectively at scale. If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you. This position offers the chance to collaborate closely with AI researchers and engineers to create cutting-edge AI solutions. Join us in shaping the future at Together AI!

Responsibilities
  • Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale.
  • Develop and optimize runtime inference services for large-scale AI applications.
  • Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
  • Conduct design and code reviews to ensure high standards of quality.
  • Create services, tools, and developer documentation to support the inference engine.
  • Implement robust and fault-tolerant systems for data ingestion and processing.
Requirements
  • 3+ years of experience writing high-performance, well-tested, production-quality code.
  • Proficiency with Python and PyTorch.
  • Demonstrated experience in building high performance libraries and tooling.
  • Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale.
  • Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum
  • Preferred: Knowledge of AI inference techniques such as speculative decoding.
  • Preferred: Knowledge of CUDA/Triton programming.
  • Nice to have: Knowledge of Rust, Cython and compilers.
About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance, and other competitive benefits. The US base salary range for this full-time position is $200,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level, and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunities to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Together AI San Francisco, California, USA Office

584 Castro St, #2050, San Francisco, California , United States, 94114

Similar Jobs

25 Days Ago
Hybrid
Palo Alto, CA, USA
178K-313K Annually
Senior level
178K-313K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Design and productionize causal machine learning models, including uplift modeling and treatment-effect estimation. Analyze A/B tests and quasi-experiments, develop experimentation strategies, evaluate modeling tradeoffs, and build scalable infrastructure. Collaborate with product and engineering teams, conduct code reviews, maintain engineering standards, communicate technical insights, and mentor others.
Top Skills: A/B TestingCausal InferenceCausalmlDowhyEconmlMachine LearningNumpyPandasPythonScikit-Learn
Yesterday
In-Office
San Francisco, CA, USA
200K-320K Annually
Entry level
200K-320K Annually
Entry level
Artificial Intelligence • Big Data • Hardware • Machine Learning
Optimize machine-learning inference for latency, throughput, and cost. The role profiles pipelines, improves kernels, batching, quantization, compilation, serving, and hardware utilization, while building serving and evaluation infrastructure. It partners with research and product teams to make models affordable at scale and treats compute cost as a core performance metric.
Top Skills: C++CudaJaxMlirNsightPythonPyTorchPytorch ProfilerRustTensorrtTvm
One Month Ago
Hybrid
Palo Alto, CA, USA
195K-343K Annually
Senior level
195K-343K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Design and deliver generative AI and computer vision models and on-device inference for consumer AR experiences. Prototype research advances, optimize models (compression, quantization, distillation), collaborate across teams, and productionize scalable inference pipelines for mobile and wearable platforms.
Top Skills: ArComputer VisionCpu OptimizationDeep LearningDiffusion ModelsGenerative ModelingGpu OptimizationJaxMl Inference PipelinesMlxModel CompressionModel DistillationNeural NetworksNpu OptimizationPyTorchQuantizationScikit-LearnTensorFlow

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account