Hippocratic AI Logo

Hippocratic AI

Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Reposted 29 Days Ago
Be an Early Applicant
In-Office
Menlo Park, CA, USA
Senior level
In-Office
Menlo Park, CA, USA
Senior level
Lead end-to-end post-training for LLMs using reinforcement learning and on-policy distillation to improve clinical reasoning, safety, and alignment. Design RL/OPD methods, build reward models and verifiers, create healthcare conversational environments with synthetic data, automate post-training loops, run experiments, and collaborate with research, engineering, and clinical teams to deploy models at scale.
The summary above was generated by AI
Role Mission

As HAI's LLM Post-Training Applied Scientist, you will own the reinforcement learning and on-policy distillation pipeline that transforms raw model capability into reliable, safe clinical behavior. Your post-training methods will directly determine how our AI agents reason through complex clinical scenarios, handle safety-critical decisions, and ultimately impact millions of patient interactions. This role exists because post-training is where capability becomes trustworthiness—and in healthcare, that's everything.

What You Will Accomplish

Own your first major outcome: By day 90, you will have shipped a post-training improvement that meaningfully advances model performance on a critical clinical capability (clinical reasoning, safety alignment, or task completion), evaluated the gains rigorously, and contributed that learning to our post-training roadmap.

Drive lasting impact: At 12 months, you will have designed and shipped multiple post-training methods that measurably improve our models' clinical safety and reasoning, built reusable infrastructure (reward models, verifiers, evaluation frameworks) that accelerate future post-training work, published your research or contributed to HAI's intellectual property, and directly shaped how our deployed models behave in production healthcare environments.

The Team

You'll work alongside ML researchers, engineers, clinicians, and safety experts who are obsessed with building trustworthy AI. This is a highly technical team that values rigor, collaboration across disciplines, and solving the hardest problems in AI safety and alignment. You'll have direct influence on model architecture and training decisions.

What You'll Do
  • Design and implement RL and OPD post-training methods including RLHF, RLVR, on-policy distillation, and novel approaches tailored to healthcare AI—selecting the right methods for different clinical reasoning and safety challenges

  • Build and evaluate reward models, verifiers, and LLM-as-judge pipelines that provide reliable training signals for post-training, ensuring they capture what truly matters in clinical contexts (accuracy, safety, patient experience)

  • Develop conversational AI environments and simulations for healthcare RL training—creating synthetic clinical scenarios and datasets that enable safe, scalable post-training without relying solely on human feedback

  • Automate post-training research loops using agents and tooling to systematically explore hyperparameters, methods, and data strategies—turning post-training into a scientific, reproducible process

  • Run rigorous experiments and analysis to understand what drives post-training gains, isolate the contributions of different components, and build intuition about what works in healthcare contexts

  • Collaborate with research, engineering, and clinical teams to translate clinical requirements into post-training objectives, validate improvements against real-world metrics, and scale successful methods to production

Location Requirement

We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Palo Alto office five days a week, unless otherwise specified..

Applied Scientist, Reinforcement Learning

Required Qualifications

  • Master's degree in Computer Science, Machine Learning, or a related field

  • 5+ years of professional experience in NLP, LLM training, or reinforcement learning

  • 2+ years of hands-on experience with RL for LLM post-training

  • Proficiency in Python and PyTorch for large-scale training

  • Demonstrated experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods

  • Experience training or fine-tuning models at scale (50B+ parameters)

Preferred Qualifications

  • Publications at top-tier ML venues (NeurIPS, ICML, ICLR, ACL, EMNLP)

  • Healthcare or regulated domain experience

  • Experience with distributed training frameworks (FSDP, DeepSpeed, vLLM)

  • Familiarity with safety alignment and interpretability research

Our comprehensive compensation package is designed to reward your expertise and includes both a competitive base salary and valuable stock options. Individual offers are determined based on a variety of factors, including your professional experience, core competencies, and geographic location.

Why Join Hippocratic AI

Reinvent healthcare with AI that puts safety first. We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact [email protected].

Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information during the hiring process.

HQ

Hippocratic AI Palo Alto, California, USA Office

167 Hamilton Ave, 3rd Floor, Palo Alto, California, United States, 94301

Similar Jobs

10 Minutes Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
200K-240K Annually
Senior level
200K-240K Annually
Senior level
Artificial Intelligence • Cloud • Software
Partners with Sales and RevOps leadership on forecasting, budgeting, headcount planning, quota capacity, commissions, bookings, and productivity. Owns Sales financial planning and business reviews, translates variances into recommendations, and builds scalable reporting, dashboards, and processes to improve forecast accuracy and operational efficiency.
Top Skills: ChatgptClaudeExcelMicrosoft Copilot
59 Minutes Ago
Hybrid
Santa Clara, CA, USA
191K-334K Annually
Senior level
191K-334K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads multiple engineering teams building scalable platform services, frameworks, infrastructure, and developer capabilities. Defines technical roadmaps with product, architecture, infrastructure, and AI teams; drives secure, reliable, cloud-native solutions and operational excellence. Responsibilities include hiring, coaching, succession planning, technical direction, platform modernization, performance optimization, and cross-functional initiative delivery. The role also advances AI-enabled engineering practices and supports ServiceNow’s AI platform strategy.
Top Skills: AgileAi/MlCloud-Native ArchitectureDevOpsDistributed SystemsEvent-Driven ArchitecturesIntelligent AutomationJavaObservabilitySreTelemetry Systems
An Hour Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
160K-270K Annually
Senior level
160K-270K Annually
Senior level
Legal Tech • Software • Generative AI
Own Eve’s messaging house, positioning, competitive narrative, first-call deck, keynote and event content. Partner with executives, Product, Sales, Enablement, Demand Gen, and Content to ensure consistent, differentiated messaging. Gather market feedback from customers, prospects, and sales calls, translating technical AI capabilities into clear value for plaintiff attorneys. The role requires exceptional writing, executive presence, strategic messaging expertise, and 7+ years of B2B SaaS or legal technology product marketing experience.
Top Skills: Artificial IntelligenceB2B SaasLegal Technology

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account