Cursor Logo

Cursor

Software Engineer, ML Platform

Posted 18 Days Ago
In-Office
San Francisco, CA, USA
Entry level
In-Office
San Francisco, CA, USA
Entry level
Build and operate ML platform infrastructure supporting researchers and product engineers. Responsibilities include designing distributed systems, data pipelines, observability tools, scheduling and orchestration infrastructure, and GPU fleet systems. The role partners closely with ML research, owns reliability and performance, improves developer experience, and ships platform primitives in a high-ownership environment.
The summary above was generated by AI

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the role

As a Software Engineer on ML Platform at SpaceXAI, you'll build the infrastructure that turns real product usage into better models — and keeps research moving fast on large GPU fleets. ML Platform is organized into four teams. Depending on your background, you may join any of them:

  • Telemetry — Own the collection and serving path that turns real product use into a record research can trust; without slowing the product, and under a small, explicit policy. Client-side or high-volume ingestion experience is a plus.

  • ML Data Platform — Build the shared environments and pipeline substrate researchers extend, so new experiments don’t fork their own stack.

  • Observability — Make it easy for researchers to start, watch, and debug their own runs.

  • ML DevX and Systems — Shorten the path from idea to a trusted run on the research fleet.

We're looking for strong distributed-systems and infrastructure engineers who want to sit next to research and ship platform primitives that move the product.

We're in-person with cozy offices in North Beach, San Francisco, Palo Alto, and Manhattan, New York, complete with well-stocked libraries.

What you’ll do
  • Design, build, and operate core platform systems used daily by ML researchers and product engineers

  • Partner closely with research to turn recurring pain into durable infrastructure

  • Own reliability, performance, and developer experience for the systems in your lane

  • Ship iteratively in a flat, high-ownership environment. Measure impact, then raise the bar

You may be a fit if
  • You have a strong background in systems / infrastructure software engineering and enjoy building platforms other engineers depend on

  • You've owned production distributed systems at meaningful scale (ingestion, data pipelines, scheduling/orchestration, or similar)

  • You're comfortable across Linux, cloud and/or bare metal, and modern orchestration (Kubernetes, Ray, or equivalent)

  • You like working closely with ML researchers and product engineers

  • You thrive where ownership is high and the feedback loop is short

Especially strong backgrounds by team
  • Telemetry: event ingestion, product analytics pipelines, OpenTelemetry / tracing, reliable data APIs

  • Product Data Platform: data frameworks, Spark / Flink / Ray, ML dataset and training-data infrastructure

  • Observability: experiment / run monitoring, debug and eval tooling, agent-friendly observability UX

  • ML DevX and Systems: GPU / cluster scheduling, job queues, node health, research compute developer experience

Applying

If there appears to be a fit, we'll reach out to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.

Cursor San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

28 Days Ago
Hybrid
Palo Alto, CA, USA
Mid level
Mid level
Financial Services
Build and operate scalable machine learning training platforms and pipelines across AWS and other cloud environments. Optimize GPU-based single-node and distributed training for performance, cost, stability, and reproducibility. Develop Kubernetes infrastructure, GenAI and LLM fine-tuning workflows, observability, CI/CD automation, governance controls, and developer self-service tools. Partner with data engineering and platform teams while applying secure, responsible AI-assisted development practices.
Top Skills: AirflowArgo WorkflowsAWSCi/CdCloudwatchDaskDdpDeepspeedEc2EcrEksFsdpGpu ComputingIamKubernetesParquetPythonPyTorchRayS3SparkTensorFlowVpcWebdataset
24 Days Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
160K-240K Annually
Mid level
160K-240K Annually
Mid level
Fintech • HR Tech
Build and scale Gusto's ML/AI platform: design MLOps solutions, automated pipelines, model development frameworks, deployment, monitoring, CI/CD, and observability; collaborate with ML engineers and apply AI tools and best practices to production ML systems.
Top Skills: Ai FrameworksAi-Assisted Development ToolsAWSCi/CdFeature StoreJavaMl ObservabilityPythonRuby
One Month Ago
Hybrid
Palo Alto, CA, USA
Senior level
Senior level
Financial Services
Lead design, build, and operate scalable ML training platform on cloud. Optimize GPU single-node and distributed training, manage Kubernetes-based infrastructure, enable GenAI/LLM fine-tuning workflows, implement observability, improve developer CI/CD and automation, and enforce governance, security, and responsible AI practices.
Top Skills: Aws IamAws S3CloudwatchDdpDeepspeedDockerEc2EcrEksFsdpGpuKubernetesPythonPyTorchTensorFlowVpc

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account