Hippocratic AI Logo

Hippocratic AI

Staff Site Reliability Engineer

Posted 5 Days Ago
Be an Early Applicant
In-Office
Menlo Park, CA, USA
Expert/Leader
In-Office
Menlo Park, CA, USA
Expert/Leader
Design and operate a GPU management and scheduling platform for approximately 30 AI models across heterogeneous hardware. Build metrics pipelines, admission control, autoscaling, cloud orchestration, infrastructure automation, deployment pipelines, and monitoring systems. Develop production software in Python and Go, operate secure fault-tolerant cloud infrastructure, enforce healthcare security and compliance policies, troubleshoot complex issues, and mentor engineers.
The summary above was generated by AI
About the Role

We're looking for a Staff Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.

We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it — collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts.

This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.

What You'll Do
  • Design and build our GPU management and scheduling platform — the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardware

  • Build the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisions

  • Implement admission control to protect capacity — deciding when to accept, queue, or shed inference requests so we operate within fleet limits

  • Build autoscaling that adjusts the number of model replicas in response to real-time demand and utilization

  • Develop cloud orchestration systems and operators in Python and Go to manage the model fleet

  • Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure

  • Design and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software

  • Stand up and maintain monitoring, logging, and alerting that keep the platform reliable and performant

  • Develop and enforce security and compliance policies appropriate to a healthcare AI platform

  • Partner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issues

  • Mentor engineers and raise the technical bar across the team

What You Bring

Must-Have

  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering

  • Computer Science Degree Required from a top CS program.

  • Strong software engineering fundamentals — you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools

  • Experience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control

  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)

  • Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)

  • Strong knowledge of containerization and orchestration (Docker, Kubernetes)

  • Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)

  • Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)

  • Excellent problem-solving skills and the ability to work both independently and collaboratively

  • Strong communication and interpersonal skills

Nice-to-Have

  • Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators

  • Familiarity with ML inference serving and model deployment (e.g. Triton, KServe, Ray Serve, or similar)

  • Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)

  • Experience implementing HIPAA and SOC 2 compliance

  • Experience operating in an HPC environment

  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field

Join our team at Hippocratic AI and help shape the future of clinically safe, production-grade AI systems.

Why Join Hippocratic AI

Reinvent healthcare with AI that puts safety first. We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact [email protected].

Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information during the hiring process.

HQ

Hippocratic AI Palo Alto, California, USA Office

167 Hamilton Ave, 3rd Floor, Palo Alto, California, United States, 94301

Similar Jobs

Yesterday
Hybrid
150K-262K Annually
Senior level
150K-262K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Maintain and improve the reliability, scalability, performance, and operability of ServiceNow’s federal cloud infrastructure. Resolve infrastructure issues, automate repetitive work, develop monitoring solutions, reduce incidents and MTTR, and lead reliability initiatives across hardware, networking, systems, applications, and cloud technologies. This third-shift role supports government customers on a Sunday-Wednesday, 10-hour schedule with no on-call rotation.
Top Skills: AgileAutomationAWSAzureCi/CdDevOpsIp AddressingJavaScriptLinuxMonitoringMySQLNetworkingObservabilityPostgresPythonRoutingRubyScriptingServicenow Platform
13 Days Ago
Hybrid
211K-263K Annually
Expert/Leader
211K-263K Annually
Expert/Leader
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, incident management, capacity planning, disaster recovery, and security engineering for Crunchyroll’s cloud-native data platforms. Design and operate Kubernetes and GCP infrastructure, implement Infrastructure as Code, improve SLOs and operational excellence, remediate vulnerabilities, strengthen cloud and container security, and mentor engineers across cross-functional teams.
Top Skills: Ci/CdDatadogDockerGCPGoGrafanaIamInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonSecrets ManagementShellSsdlcTerraformZero Trust
13 Days Ago
Hybrid
San Francisco, CA, USA
233K-292K Annually
Senior level
233K-292K Annually
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Design and operate Kubernetes and GCP systems, establish SRE practices including SLIs, SLOs, and error budgets, manage incidents and vulnerabilities, strengthen cloud security, and mentor engineers across teams.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonSecrets ManagementSecure Software Development LifecycleShellTerraform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account