Hippocratic AI Logo

Hippocratic AI

Machine Learning Engineer

Reposted 2 Days Ago
In-Office
Menlo Park, CA, USA
Entry level
In-Office
Menlo Park, CA, USA
Entry level
Build reliable training, evaluation, deployment, and data-feedback pipelines for a recursively self-improving machine learning system. Design reward and feedback signals, mitigate reward hacking and distribution drift, develop evaluation harnesses, debug model regressions, stabilize non-stationary learning loops, and ship production ML systems. The role requires strong Python and machine learning engineering fundamentals, reinforcement learning knowledge, production experience, and familiarity with feedback-driven systems such as RLHF, RLAIF, active learning, or agent evaluation.
The summary above was generated by AI
About the role

We are building a recursive self-improvement system — a machine learning system that iteratively improves itself through feedback, evaluation, and automated learning loops. You will help build the engineering pipeline that keeps these loops fast, reliable, and trustworthy: the training and evaluation pipelines, the reward and feedback signals, and the safeguards that prevent a self-improving system from silently degrading or gaming its objectives.

This is an engineering-first role with deep reinforcement learning requirements. You should be equally comfortable writing robust production ML code and reasoning about reward design, credit assignment, and why feedback-driven systems become unstable.

What you'll do
  • Build and maintain the training, evaluation, and deployment loops at the core of the self-improvement system, with a strong emphasis on reproducibility and reliability.

  • Design and implement reward and feedback signals; investigate and mitigate reward hacking, specification gaming, and distribution drift.

  • Build evaluation harnesses and metrics before models — because a self-improving system is only as safe as its measurement of “better.”

  • Own data pipelines and automated data flywheels that feed the learning loop.

  • Debug subtle model-quality regressions and stabilize training and feedback loops that go non-stationary.

  • Collaborate with research and product to turn methods into robust, shippable systems.

What we're looking for

Must-Haves:

Strong MLE fundamentals (non-negotiable)

  • Excellent Python and clean, well-tested ML training code.

  • Solid grasp of data pipelines, distributed / large-scale training, and experiment tracking.

  • The instinct and skill to debug why a model silently got worse — not just why it crashed.

Hands-on experience with a feedback or learning loop (at least one).

  • Built or owned part of a feedback loop — a reward model, an evaluation harness, or the data pipeline for an RLHF/RLAIF or active-learning system.

  • Ran a retraining or continual-learning pipeline where a model consumed its own predictions or production data (e.g. ranking, recommendations, fraud, spam).

  • Fine-tuned LLMs with human or AI feedback, or built agentic evaluation harnesses.

Reinforcement learning foundations and curiosity.

  • Working knowledge of reward modeling, on-policy vs. off-policy tradeoffs, and credit assignment (does not need to be a research-level RL expert).

  • Has seen — or can reason clearly about — feedback-system failure modes: reward hacking, specification gaming, feedback loops amplifying errors.

  • Comfortable evaluating non-stationary systems (systems whose behavior and data distribution change over time).

Systems and evaluation instinct.

  • Builds the eval before the model; treats measurement as a first-class deliverable.

  • Has shipped an ML system into production and kept it healthy over time.

Strongest signal

Bonus, not required:

The ideal candidate has built or shipped a full system that improved from its own outputs or feedback end to end. This is rare at this level, so treat it as a standout differentiator rather than a filter. Examples:

  • RLHF / RLAIF pipelines

  • Self-play systems

  • Active-learning loops

  • Automated data flywheels

  • Agentic evaluation harnesses

Nice to Have:

  • PhD or MS in RL / ML paired with real production experience (either the science or the engineering half alone is fine if the other is strong).

  • Experience at a lab or company doing RLHF, agents, or large-scale ML infrastructure.

  • Familiarity with LLM fine-tuning, evaluation frameworks, or agent orchestration.

Why Join Hippocratic AI

Reinvent healthcare with AI that puts safety first. We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact [email protected].

Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information during the hiring process.

HQ

Hippocratic AI Palo Alto, California, USA Office

167 Hamilton Ave, 3rd Floor, Palo Alto, California, United States, 94301

Similar Jobs

Yesterday
Hybrid
2 Locations
209K-286K Annually
Senior level
209K-286K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Build, deploy, scale, and operate production machine learning models and platforms. Responsibilities include developing ML components, distributed data pipelines, cloud and Kubernetes infrastructure, model monitoring and retraining, CI/CD automation, responsible AI governance, and collaboration with product, data science, and engineering teams.
Top Skills: AgileSparkAWSAzureC++Ci/CdGCPGoJavaKubernetesNumpyPandasPythonPyTorchRayScalaScikit-LearnTensorFlow
Yesterday
Hybrid
San Francisco, CA, USA
209K-286K Annually
Senior level
209K-286K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Build, deploy, scale, and maintain production machine learning models and platforms. Responsibilities include designing ML systems, developing data pipelines, operating distributed systems and cloud infrastructure, automating testing and deployment, monitoring and retraining models, and applying responsible AI practices. The role collaborates with Product and Data Science teams and uses Python, Java, Scala, or related languages.
Top Skills: AgileSparkAWSAzureC++Ci/CdGCPGoJavaKubernetesNumpyPandasPythonPyTorchRayScikit-LearnTensorFlow
2 Days Ago
Hybrid
2 Locations
230K-286K Annually
Senior level
230K-286K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Build, deploy, scale, and maintain machine learning models and platforms in production. Develop optimized data pipelines, distributed ML systems, cloud-based architectures, and automated CI/CD workflows. Collaborate with product and data science teams, monitor and retrain models, apply responsible AI practices, and ensure secure, governed software delivery. The role requires extensive experience with ML frameworks, distributed systems, cloud services, Kubernetes, and production machine learning operations.
Top Skills: SparkAWSAzureC++Ci/CdDockerGCPGoJavaKubernetesNumpyPandasPythonPyTorchRayScalaScikit-LearnTensorFlow

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account