SPREEAI Logo

SPREEAI

MLOps Engineer

Posted 2 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
145K-180K Annually
Entry level
In-Office
San Francisco, CA, USA
145K-180K Annually
Entry level
Own the end-to-end ML lifecycle platform, including training-as-a-service, experiment tracking, model CI/CD, data versioning, monitoring, and cost governance. Build infrastructure for multi-GPU training, automated evaluation gates, canary releases, A/B testing, and checkpoint management. Partner with ML Scientists to improve reproducibility and workflow efficiency, integrate external model providers, and help guide platform direction while mentoring engineers.
The summary above was generated by AI

SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences. Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.

About the team

AI Platform scales SPREEAI's infrastructure: productionizing ML model checkpoints, running the API services behind the Partner Portal, optimizing inference serving, and giving ML Scientists training-as-a-service so they can iterate without managing infrastructure themselves.

About the role

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

What you'll do
  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale
  • Partner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions
  • Evaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform
What you'll bring
  • Experience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)
  • Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment
You'll thrive here if

You treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident. You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.

Similar Jobs

Yesterday
In-Office
San Francisco, CA, USA
66K-165K Annually
Mid level
66K-165K Annually
Mid level
Healthtech • Biotech • Pharmaceutical
Build and operate MLOps infrastructure for scientific discovery platforms, including model deployment pipelines, registries, monitoring, APIs, cloud-native serving, CI/CD, and observability. The role operationalizes computational biology, cheminformatics, and related scientific models for researchers and AI agents, ensuring scalable performance, versioning, reliability, uncertainty metrics, and robust integration with data pipelines.
Top Skills: ApptainerAWSAzureCC++Ci/CdCudaDvcGCPGrpcKubeflowKubernetesLangchainMcpMlflowPythonPyTorchRest ApisScikit-LearnSingularityTensorFlowWeights & Biases
7 Days Ago
In-Office
San Francisco, CA, USA
80K-210K Annually
Senior level
80K-210K Annually
Senior level
Artificial Intelligence • Logistics • Software • Defense
Build and operate production ML infrastructure across cloud, on-premises, GPU, edge, and disconnected environments. Own model training and serving, CI/CD, reproducibility, evaluation, observability, data pipelines, retrieval systems, and secure IL5/IL6 deployments. Support accreditation and classified environments while improving reliability, quality, and performance of LLM and ML systems.
Top Skills: Amazon EksAmazon SagemakerArgocdAWSAzureCi/CdCmmcDockerFips 140-3GitopsGpu InfrastructureInfrastructure As CodeJetstreamKubernetesNatsNist Sp 800-171Nist Sp 800-53PgvectorPostgresPythonRmf/EmassTensorrt-LlmVllm
10 Days Ago
In-Office or Remote
California, USA
152K-230K Annually
Senior level
152K-230K Annually
Senior level
Software • Travel
Build and maintain internal platforms and tooling for AI engineers, including React and Next.js dashboards, Python backend services, REST APIs, model deployment automation, observability integrations, and ML pipeline testing. Partner with engineering and ML teams to improve developer experience, production AI reliability, and self-service workflows while documenting systems and runbooks.
Top Skills: AWSAzureDatabasesDatadogDockerGCPGrafanaInfrastructure As Code (Iac)Job Queuing SystemsKubernetesLangsmithNext.JsPrometheusPythonReactRest ApisTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account