Together AI Logo

Together AI

Senior AI Infrastructure Engineer

Sorry, this job was removed at 09:09 p.m. (PST) on Monday, Jun 02, 2025
Be an Early Applicant
In-Office
San Francisco, CA, USA
In-Office
San Francisco, CA, USA

Similar Jobs

17 Days Ago
Hybrid
Palo Alto, CA, USA
Senior level
Senior level
Financial Services
Designs, builds, and operates secure, scalable cloud and GPU infrastructure platforms for enterprise AI/ML workloads. Leads architecture, production coding, Kubernetes and container operations, CI/CD, infrastructure automation, performance optimization, and reliability efforts. Partners with AI/ML and platform teams to support distributed multi-GPU training and inference. Provides technical leadership while advancing responsible AI-assisted engineering, secure SDLC practices, automation, and operational excellence.
Top Skills: BcmC#Ci/CdCloud InfrastructureCudaDistributed SystemsDockerGoGpu InfrastructureJavaKubernetesLinuxMicroservicesMlflowNvidia DcgmNvidia DriversPythonRay.IoSlurm
2 Days Ago
In-Office
Palo Alto, CA, USA
Entry level
Entry level
Artificial Intelligence • Machine Learning • Software • Generative AI
Build and optimize foundational AI infrastructure for large-scale training, reinforcement learning, rollout, and inference workloads. Responsibilities include profiling bottlenecks, optimizing GPU and distributed execution, developing scheduling and load-balancing systems, modeling performance and capacity, improving fault tolerance, and creating benchmarks and observability tools. The role partners closely with AI researchers and systems engineers to improve throughput, reliability, utilization, latency, and cost across complex GPU clusters.
Top Skills: C++Cluster SchedulersContainer RuntimesCudaDeepspeedFsdpGoInfinibandJaxKubernetesLinuxMegatronNcclNvlinkPythonPyTorchRayRcclRdmaRoceRustSglangStorage SystemsTensorrt-LlmTorch.DistributedTritonVllm
5 Days Ago
In-Office
Palo Alto, CA, USA
165K-265K Annually
Senior level
165K-265K Annually
Senior level
Aerospace • Other
Designs, deploys, operates, and scales GPU/CPU infrastructure supporting Starshield national security missions. Responsibilities include managing Kubernetes and AI clusters, bare-metal and virtualized GPU platforms, operating systems, databases, monitoring, distributed storage, automation, and high-availability services. The role collaborates with AI engineers, improves service lifecycles, mentors junior engineers, and leads technical excellence. Requires Linux, Kubernetes, infrastructure automation, scripting, software development, and the ability to obtain a Top Secret clearance.
Top Skills: AnsibleBashBazelC++Distributed DatabasesDistributed StorageGoKubernetesLinuxMakefilesMonitoring And AlertingNvidia Gpu Deployment StacksOci ContainersPythonTcp/IpTerraformVirtualization
About the Role

As a Senior AI Infrastructure Engineer, you will be responsible for building the next generation, highly available, global, multi-cloud PaaS platform with open-source technologies to enable and accelerate Together AI’s rapid growth.

This system spans many diverse environments (Kubernetes, VMs, bare metal compute, and edge deployments) and provides a cohesive and reliable abstraction for running AI workloads in them. You will get to be a technology thought leader, evangelize new, cutting-edge technologies, and solve complex problems.

To be successful, you’ll need to be deeply technical and possess excellent communication, collaboration, and diplomacy skills. You have experience practicing infrastructure-as-code, including using tools like Terraform and Ansible. You have strong software development fundamentals and skills. In addition, you have strong systems knowledge and troubleshooting abilities.

Requirements
  • 5+ years of professional software development experience and proficiency in at least one backend programming language (Golang desired)
  • Demonstrated experience with high performance or distributed cloud microservices architectures and ideally experience building them in operation at a global scale using multiple cloud providers such as AWS, Azure, or GCP
  • Excellent understanding of low level operating systems concepts including multi-threading, memory management, networking and storage, performance, and scale
  • Pragmatic, methodical, well-organized, detail-oriented, and self-starting
  • Experience with Kubernetes and containerization, VPNs, AI workloads, and blockchain based protocols a plus
  • GPU programming, NCCL, CUDA knowledge a plus
  • Experience with Pytorch or Tensorflow a plus
  • 5+ years experience writing high-performance, well-tested, production quality code
Responsibilities
  • Perform architecture and research work for decentralized AI workloads
  • Work on the core, open-source Together AI platform
  • Create services, tools, and developer documentation
  • Create testing frameworks for robustness and fault-tolerance
About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Together AI San Francisco, California, USA Office

584 Castro St, #2050, San Francisco, California , United States, 94114

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account