Raydar Logo

Raydar

Software Engineer - ML Infrastructure

Posted 18 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
250K-300K Annually
Senior level
In-Office
San Francisco, CA, USA
250K-300K Annually
Senior level
Own machine learning infrastructure spanning distributed training, reinforcement learning, data pipelines, model serving, deployment, monitoring, and rollout tooling. Partner with researchers to productionize experimental workflows for clinical AI applications. The role also involves GPU scheduling, autoscaling, reproducible environments, online-learning loops, checkpointing, evaluation, and technical leadership or mentoring.
The summary above was generated by AI

About the company

Our client is a healthcare technology company developing AI-enabled imaging tools. Researchers and engineers work together on production machine learning systems for clinical applications.

The role

Raydar is recruiting for this role on behalf of our client. Own ML infrastructure across distributed training, reinforcement learning, and production serving. You will establish engineering practices and help researchers translate experiments into reliable systems.

What you'll do

- Build distributed training infrastructure with parallelism and checkpointing.

- Develop reinforcement-learning infrastructure for rollout generation, reward models, and experience collection.

- Partner with researchers to turn experimental workflows into production systems.

- Build data loading and preprocessing pipelines for multimodal datasets.

- Improve model serving, canary deployment, monitoring, and rollout tooling.


Requirements

What we're looking for

- 6+ years of ML infrastructure or distributed systems experience.

- Strong Python and hands-on production ML serving experience.

- Kubernetes and Docker experience with GPU scheduling, autoscaling, and reproducible environments.

- Distributed training expertise in PyTorch or JAX, including FSDP, DeepSpeed, or comparable approaches.

- Experience with RL or online-learning loops, logging, checkpointing, and evaluation.

- Breadth across the ML infrastructure stack plus technical leadership or mentoring experience.

Bonus points

- Startup experience.

- A/B testing and production ML experimentation platforms.

- A computer science or other STEM degree.


Benefits

Compensation and benefits

- Base salary: USD 250,000 to 300,000 per year.

- Equity: Competitive equity.

Location and work model

- San Francisco, California, United States.

- Five days per week onsite.

- Visa transfers and new visa sponsorships are supported.

Similar Jobs

15 Days Ago
Hybrid
Palo Alto, CA, USA
178K-313K Annually
Senior level
178K-313K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Build and optimize Snap’s machine learning infrastructure, including scalable training, evaluation, inference, feature serving, data management, vector search, and model deployment systems. Improve reliability, performance, and cost efficiency for large-scale ML workloads while collaborating with ML engineers. The role requires strong programming, distributed systems, big data, cloud infrastructure, and production machine learning experience.
Top Skills: Ai Model InferenceBig Data ProcessingC++Caffe2Cloud InfrastructureDistributed SystemsFlinkJavaMachine Learning InfrastructurePythonPyTorchRayScalaScikit-LearnSparkSpark MlTensorFlowVector Search
7 Days Ago
In-Office or Remote
2 Locations
245K-429K Annually
Senior level
245K-429K Annually
Senior level
Social Media
Set technical vision for Pinterest’s Product ML Infrastructure, leading distributed model training, fine-tuning, evaluation, and high-scale CPU/GPU inference. Improve GPU efficiency, data loading, execution, kernels, memory, compilation, quantization, scheduling, and capacity. Build reliable, observable training and serving platforms, drive architecture decisions and migrations, partner with AI/ML teams, mentor senior engineers, and establish operational standards across the organization.
Top Skills: Ads RankingC++CompilationDistributed Machine Learning SystemsFeature PlatformsGpuGpu KernelsJavaModel TrainingOnline InferencePythonQuantizationRecommender SystemsRetrieval Systems
One Month Ago
In-Office
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Information Technology • Robotics • Software
Design, own, and scale training and inference infrastructure for a robotic fleet. Build data pipelines, optimize distributed training and GPU utilization, speed experiment launch and reproducibility, and contribute to core training code.
Top Skills: GpuPythonPyTorchTensorFlow

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account