Eli Lilly and Company Logo

Eli Lilly and Company

ML Ops Engineer

Posted 3 Days Ago
In-Office
South San Francisco, CA, USA
147K-268K Annually
Mid level
In-Office
South San Francisco, CA, USA
147K-268K Annually
Mid level
Build and operate scalable machine learning platforms supporting drug discovery. Responsibilities include model deployment, monitoring, retraining, reliability, inference optimization, GPU resource management, infrastructure automation, CI/CD, observability, performance tuning, and production integration. Collaborate with research scientists, data scientists, AI engineers, and infrastructure teams to deliver reliable, reproducible AI capabilities.
The summary above was generated by AI

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. 


Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley!

Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health.  Are you up for the challenge?  If so, join us! 

About the Lilly and NVIDIA Partnership

Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.

What You’ll Be Doing

As an ML Ops Engineer, you build and operate the platforms that run the end-to-end machine learning lifecycle. You enable reliable model deployment, operation, monitoring, retraining, and reproducibility at scale. You optimize infrastructure and GPU resources to support research and discovery workloads. You will work closely with engineering and scientific teams to deliver production-ready AI capabilities.

How You’ll Succeed

  • Lead the operational lifecycle of ML models, including deployment, monitoring, and ongoing reliability.

  • Operate and optimize large-scale inference platforms that support scientific discovery and AI workloads.

  • Ensure models can be deployed, scaled, monitored, and maintained in production environments.

  • Test, refine, and improve model accuracy.

  • Work with data scientists, business analysts and partners to integrate ML models into broader strategies.

  • Automate the platform with infrastructure-as-code and CI/CD, and document it well enough that someone else can operate it.

What You Should Bring

  • Strong Python skills and experience working with machine learning frameworks such as PyTorch, JAX, or TensorFlow.

  • Experience deploying, operating, and scaling production machine learning platforms, including model serving, monitoring, and large-scale inference workloads.

  • Experience with MLOps platforms and tools like MLflow, Weights & Biases, KServe, or similar technologies.

  • Proficiency with containerization, orchestration, and distributed compute environments (Docker, Kubernetes, Slurm, Ray).

  • Experience operating large-scale AI platforms that deploy, host, and optimize machine learning models for production use, using technologies such as Triton, vLLM, or TensorRT-LLM.

  • Experience with infrastructure automation and CI/CD practices using tools such as Terraform, Ansible, GitHub Actions, or related.

  • Experience supporting cloud platforms (AWS, Azure, or GCP) and on-premises GPU infrastructure.

  • Knowledge of observability and operational monitoring, including metrics, logging, tracing, and performance tuning.

  • Ability to identify and address system, infrastructure, and model performance issues through automation and continuous improvement.

  • Ability to collaborate effectively with research scientists, AI engineers, and infrastructure teams in a fast-paced environment.

Your Basic Qualifications

  • Bachelor’s in Computer Science, Engineering, Statistics, Mathematics, or a related technical field

  • 4+ years of experience in machine learning engineering, ML Ops, or platform engineering.

Location & Work Flexibility
This role is based at our Silicon Valley Hub. We offer a flexible hybrid work model, with three days onsite and two days working remotely each week, supporting both collaboration and work‑life balance.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.


Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.


Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).


Actual compensation will depend on a candidate’s education, experience, skills, and geographic location.  The anticipated wage for this position is

$147,000 - $268,400

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Similar Jobs

19 Days Ago
In-Office
San Mateo, CA, USA
280K-420K Annually
Senior level
280K-420K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead the architecture and implementation of a Kubernetes-native AI Factory platform for distributed training, simulation, evaluation, inference, and deployment. Build GPU infrastructure, scheduling, storage, networking, observability, model lifecycle capabilities, and self-service developer workflows across cloud, on-premises, sovereign, and air-gapped environments. Partner with ML researchers and autonomy teams while providing technical leadership on scalable, maintainable AI infrastructure.
Top Skills: GoGrafanaHelmHugging Face TransformersKaiKubernetesLinuxOpentelemetryPrometheusPythonPyTorchRaySlurmTerraform
2 Days Ago
In-Office
San Francisco, CA, USA
80K-210K Annually
Senior level
80K-210K Annually
Senior level
Artificial Intelligence • Logistics • Software • Defense
Build and operate production ML infrastructure across cloud, on-premises, GPU, edge, and disconnected environments. Own model training and serving, CI/CD, reproducibility, evaluation, observability, data pipelines, retrieval systems, and secure IL5/IL6 deployments. Support accreditation and classified environments while improving reliability, quality, and performance of LLM and ML systems.
Top Skills: Amazon EksAmazon SagemakerArgocdAWSAzureCi/CdCmmcDockerFips 140-3GitopsGpu InfrastructureInfrastructure As CodeJetstreamKubernetesNatsNist Sp 800-171Nist Sp 800-53PgvectorPostgresPythonRmf/EmassTensorrt-LlmVllm
11 Days Ago
In-Office or Remote
10 Locations
105K-141K Annually
Mid level
105K-141K Annually
Mid level
Insurance
Operate and improve enterprise AI/ML platforms across Palantir Foundry, AWS Bedrock, and Amazon SageMaker. Responsibilities include deployment, monitoring, observability, automation, governance, model lifecycle management, incident response, root cause analysis, CI/CD, infrastructure as code, and production reliability. The role provides technical leadership and mentorship while partnering with data, infrastructure, security, architecture, product, and engineering teams to support cloud-based AI workloads and Generative AI adoption.
Top Skills: Amazon CloudwatchAmazon SagemakerAWSAws BedrockCi/CdDatabricksDatadogGenerative AiGithub ActionsGrafanaInfrastructure As CodeJavaJavaScriptJenkinsLarge Language ModelsMavenMlopsNode.jsOpentelemetryPalantir FoundryPrometheusPythonRetrieval-Augmented Generation (Rag)SplunkTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account