Handshake Logo

Handshake

Senior Software Engineer, Machine Learning Infrastructure

Reposted One Month Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
176K-220K Annually
Senior level
In-Office
San Francisco, CA, USA
176K-220K Annually
Senior level
Build and operate shared ML/AI infrastructure: data pipelines, feature stores, training, model serving and LLM platform work. Scale inference and training (GPU, batching, autoscaling), enable evaluation and post-training workflows, and partner with AI, data science, and product teams to productionize models and improve platform reliability and developer experience.
The summary above was generated by AI

About Handshake

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.

Why join Handshake now:

  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel

  • Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world’s top educational institutions

  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders

  • Build a massive, fast-growing business with billions in revenue

About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

We’re looking for a Senior Software Engineer to join our ML Infrastructure & Platform team. This team powers both Handshake’s core career marketplace and Handshake AI by building the shared infrastructure behind our production ML and AI systems.

This is an infrastructure-heavy role for an engineer who enjoys building scalable platforms at the intersection of software engineering, machine learning, and generative AI. You’ll help teams move quickly from prototype to production while building the reliable, high-performance systems that power training, evaluation, and inference across Handshake.

What You’ll Do
  • Build and operate the shared infrastructure behind production ML and AI, including data pipelines, feature stores, training, and model serving.

  • Develop and scale our LLM platform, including provider integrations, orchestration, observability, and controls for cost, latency, and reliability.

  • Build evaluation infrastructure, including LLM eval harnesses, benchmarks, and quality measurement pipelines.

  • Support post-training workflows, including fine-tuning, reinforcement learning pipelines, and supporting data infrastructure.

  • Optimize inference infrastructure for open and fine-tuned models, including GPU serving, batching, and autoscaling.

  • Partner with AI, Data Science, and Product teams to productionize new models and establish best practices for ML infrastructure across Handshake.

  • Improve the reliability, scalability, and developer experience of our ML platform.

Desired Capabilities
  • 5+ years of production software engineering experience using Python, Go, TypeScript, or similar languages.

  • Experience building and operating cloud infrastructure on AWS, GCP, or similar platforms.

  • Strong experience with Kubernetes, Docker, Terraform, CI/CD, and operating production services.

  • Hands-on experience building ML infrastructure, including model serving, training pipelines, feature stores, embeddings, or ML observability.

  • Experience with modern data platforms such as BigQuery, Airflow, Spark, Beam/Dataflow, or streaming pipelines.

  • Practical experience building production systems with LLMs or generative AI, including orchestration, provider APIs, observability, and performance optimization.

  • Strong systems design skills, sound engineering judgment, and the ability to thrive in ambiguous, fast-moving environments.

Extra Credit
  • Experience with Ray, Anyscale, KubeRay, Ray Serve, vLLM, Triton, PyTorch, or GPU-backed inference and training.

  • Experience designing LLM evaluation frameworks, benchmarking systems, or quality regression testing.

  • Experience with Vertex AI, Bigtable, Redis, or feature platform infrastructure.

  • Experience with post-training techniques such as fine-tuning, RLHF, reinforcement learning, or reward modeling.

  • Experience building agentic systems, MCP integrations, tool use, memory systems, or voice AI applications.

Perks

Handshake delivers benefits that help you feel supported—and thrive at work and in life.

The below benefits are for full-time US employees.

🎯 Ownership: Equity in a fast-growing company

💰 Financial Wellness: 401(k) match, competitive compensation, financial coaching

🍼 Family Support: Paid parental leave, fertility benefits, parental coaching

💝 Wellbeing: Medical, dental, and vision, mental health support, $500 wellness stipend

📚 Growth: $2,000 learning stipend, ongoing development

💻 Remote & Office: Internet, commuting, and free lunch/gym in our SF office

🏝 Time Off: Flexible PTO, 15 holidays + 2 flex days

🤝 Connection: Team outings & referral bonuses

Explore our mission, values, and comprehensive US benefits at joinhandshake.com/careers.

HQ

Handshake San Francisco, California, USA Office

We're located right in the center of everything in the financial district of downtown San Francisco. We're just 1 block from Montgomery St Bart!

Similar Jobs

6 Days Ago
Hybrid
Palo Alto, CA, USA
Senior level
Senior level
Financial Services
Designs, builds, and operates secure, scalable cloud and GPU infrastructure platforms for enterprise AI/ML workloads. Leads architecture, production coding, Kubernetes and container operations, CI/CD, infrastructure automation, performance optimization, and reliability efforts. Partners with AI/ML and platform teams to support distributed multi-GPU training and inference. Provides technical leadership while advancing responsible AI-assisted engineering, secure SDLC practices, automation, and operational excellence.
Top Skills: BcmC#Ci/CdCloud InfrastructureCudaDistributed SystemsDockerGoGpu InfrastructureJavaKubernetesLinuxMicroservicesMlflowNvidia DcgmNvidia DriversPythonRay.IoSlurm
One Month Ago
In-Office
Mountain View, CA, USA
194K-291K Annually
Senior level
194K-291K Annually
Senior level
Artificial Intelligence • Automotive • Information Technology • Robotics
Build and operate large-scale ML training infrastructure: distributed GPU training, multi-cluster scheduling, data pipelines (batch & streaming), ML workflows, observability, alerting, and on-call/incident response to ensure reliable, cost-effective model training and releases.
Top Skills: Batch Data PipelinesC++Distributed Gpu TrainingGCPGoGpusKubernetesMulti-Cluster SchedulingNcclObservability/MonitoringPythonReinforcement LearningStreaming Data Pipelines
One Month Ago
In-Office
San Francisco, CA, USA
170K-190K Annually
Senior level
170K-190K Annually
Senior level
Artificial Intelligence • Logistics • Robotics • Software
Design, build, and operate large-scale ML infrastructure and GPU compute clusters for computer vision and multi-modal models. Own end-to-end pipelines from data ingestion and training to low-latency cloud deployment, orchestration, monitoring, and model performance evaluation. Collaborate with research and product teams to productionize models across warehouse environments.
Top Skills: AirflowC++CudaDeepspeedFlyteGpuGrpcKafkaNvidia Triton Inference ServerPythonPyTorchQuantizationRos2TemporalTensorFlowTensorflow ServingTensorrtTorchserve

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account