Tavus Logo

Tavus

Software Engineer, Infrastructure

Reposted 8 Days Ago
Remote or Hybrid
Hiring Remotely in San Francisco, CA, USA
Senior level
Remote or Hybrid
Hiring Remotely in San Francisco, CA, USA
Senior level
As a Senior Software Engineer in Infrastructure, you'll manage GPU and cloud infrastructure, ensuring stability and enhancing developer experience in a fast-paced environment.
The summary above was generated by AI
About Us

Tavus is a research lab pioneering human computing. We’re building AI Humans: a new interface that closes the gap between people and machines, free from the friction of today’s systems. Our real-time human simulation models let machines see, hear, respond, and even look real—enabling meaningful, face-to-face conversations. AI Humans combine the emotional intelligence of humans with the reach and reliability of machines, making them capable, trusted agents available 24/7, in every language, on our terms.

Imagine a therapist anyone can afford. A personal trainer that adapts to your schedule. A fleet of medical assistants that can give every patient the attention they need. With Tavus, individuals, enterprises, and developers can all build AI Humans to connect, understand, and act with empathy at scale.

We’re a Series B company backed by world-class investors including Sequoia Capital, Y Combinator, and Scale Venture Partners.

Be part of shaping a future where humans and machines truly understand each other.

The Role

We're hiring a Senior Software Engineer (Infrastructure) to own the systems behind CVI, our real-time conversational product. Every live conversation between a person and a PAL runs on infrastructure your team owns. You'll take goals like uptime, latency, and cost and chase them wherever they lead, including into backend services and product code.

What you'll own
  • CVI's inference deployments. The GPU infrastructure serving live conversations across multiple providers and regions. You'll join as an early senior member of a growing infra team, working on projects like tuning the newest GPU generations and cutting cold-start and model load times so users wait less.

  • Expanding our GPU footprint. You'll bring on new providers and regions, stand up clusters on EKS, and build the routing, scheduling, and throughput needed for fast weight loading.

  • Uptime. You'll be one of the people pushing our uptime bar higher, along with the security and SOC2 work that keeps our infrastructure trustworthy.

  • Fix what you find. When you see a problem, you have the trust and the mandate to fix it or flag it. Reworking our deploy pipeline so shipping is fast and boring is exactly the kind of thing you'd take on.

What this role has shipped
  • Multi-provider, multi-region inference infrastructure: routes live conversations across GPU providers and regions, so one provider's outage never becomes a user's problem

  • CUDA optimizations for Phoenix, our video rendering model: doubled the frame rate by tracing and optimizing hot paths with our researchers

  • Parallel conversations on a single GPU: several live conversations sharing one card, multiplying what the fleet can serve

Who you are
  • You own outcomes. You don't stop where "infrastructure" ends. If the fix lives in backend code or the CVI stack, you dive in, and you don't wait for a ticket to do it.

  • You're energized by unfamiliar problems. If the next thing that matters is standing up a training deployment you've never touched, you jump in and learn on the fly.

  • You adapt as priorities evolve. In a space moving this fast, the most important thing to build can change as we learn. When it does, you adjust course without losing momentum.

  • You care about this problem. Keeping large-scale, real-time systems fast and reliable is something you think about unprompted.

Requirements
  • Hands-on GPU inference experience. You've deployed and optimized inference workloads on GPUs and know what it takes to build reliable systems on top of GPU cloud providers.

  • Kubernetes and EKS depth, including routing and scheduling. You're comfortable designing how work gets placed across a fleet, and writing the services that make it happen.

  • Deep AWS experience. You're at home spinning up new services and turning them into simple, repeatable processes others can build on.

  • A senior track record of ownership. You've set technical direction, made decisions others built on, and carried ambiguous work over the finish line. You explain complex ideas clearly, to engineers and non-engineers alike.

Nice to have
  • Experience with GCP

  • Experience with video streaming infrastructure

  • Experience with training infrastructure or LLM serving

  • Experience with SOC2 or security compliance

If you don't check every box but this sounds like the work you want to be doing, apply anyway.

HQ

Tavus San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

5 Days Ago
Easy Apply
Remote
United States
Easy Apply
173K-255K Annually
Senior level
173K-255K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead delivery for the Batch Infrastructure team: design, build, and operate reliable, scalable compute platforms for scheduled and on-demand batch workloads. Collaborate with product/design/analytics, define technical plans, ensure availability (monitoring/on-call), set code and design standards, and mentor engineers to improve delivery and quality.
Top Skills: AirflowAWSFlinkFlyteKotlinKubernetesLuigiMySQLPrefectPythonSparkTemporal
3 Days Ago
In-Office or Remote
Santa Clara, CA, USA
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and maintain large-scale AI/ML platform and infrastructure for training, inference, fine-tuning, and Agentic AI. Develop tools, APIs, reliability metrics, and root-cause analyses across application to hardware layers while improving efficiency, resiliency, monitoring, and observability.
Top Skills: C/C++Cloud-NativeDgx CloudDynamoElkGoInfinibandJaxKubernetesLokiNcclNvidia GpusObservability PlatformsPrometheusPythonPyTorchRayRdmaScripting LanguagesTensorFlow
4 Days Ago
In-Office or Remote
2 Locations
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and maintain scalable infrastructure and CI/CD pipelines for TensorRT Edge-LLM across embedded and desktop GPU/CPU platforms. Configure and operate tools (CMake, GitLab, GitHub Actions, Kubernetes, Docker), monitor systems, automate builds/tests, address security CVEs, and collaborate with autonomy and external partners to improve development velocity and reliability.
Top Skills: AWSAzureBazelCC++CmakeDockerDrive AgxGCPGitGithub ActionsGitlabJenkinsJetpackJetson AgxKubernetesMakePerforcePythonQnxSglangTensorrtTensorrt-LlmUbuntuVllm

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account