SPREEAI Logo

SPREEAI

Principal Engineer, AI Platform & Infrastructure

Posted 2 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Expert/Leader
In-Office
San Francisco, CA, USA
Expert/Leader
Lead the design and operation of SPREEAI’s ML platform, including training workflows, model deployment, inference serving, GPU infrastructure, observability, evaluation, and production monitoring. Establish reliable research-to-production pipelines, model governance, SLOs, rollback strategies, and performance optimization across distributed systems. Partner with research, product, and engineering teams to deliver scalable, cost-efficient multimodal AI systems for real-time virtual try-on experiences.
The summary above was generated by AI

SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences. Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.

About the Role

SPREEAI is building the future of AI-powered commerce through photorealistic virtual try-on and multimodal intelligence. We bring together cutting-edge AI and real-world retail to deliver production systems that redefine how people shop online.


We are looking for a Principal Engineer to build the infrastructure, deployment pipelines, and observability systems that enable multimodal AI models to move from research prototypes to reliable, production-grade deployments powering real-time virtual try-on experiences for global retail partners.


This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability. You will partner closely with Applied Science, AI Platform, Product, and Partner Engineering to enable rapid research iteration and reliable model delivery at scale.
What You'll Own

ML Platform & Training Enablement

  • Build and operate SPREEAI’s end-to-end ML platform spanning training, evaluation, deployment, and monitoring.
  • Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems.
  • Define platform standards for model packaging, model registry, dataset lineage, experiment tracking, checkpointing, and deployment automation.


Deployment, Inference & Observability

  • Enable reliable and scalable inference deployments through standardized serving, orchestration, and monitoring frameworks.
  • Build and operate model deployment pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability.
  • Establish production SLOs for latency, availability, error rate, GPU saturation, cold-start time, cost per inference, and model quality drift.
  • Standardize and support serving infrastructure using modern inference runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, Ray Serve, TorchServe, ONNX Runtime, or equivalent systems.


GPU Infrastructure & System Efficiency

  • Design and manage GPU allocation, scheduling, and resource utilization across training and inference workloads.
  • Improve GPU utilization, throughput, latency, reliability, and cost efficiency across model lifecycle systems.
  • Design and operate model evaluation and benchmarking systems, including automated regression detection and quality gates for production releases.
  • Partner with research teams to productionize new capabilities by providing robust infrastructure, tooling, and deployment pathways.


What We're Looking For

  • 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.
  • Deep experience with Python, PyTorch, Kubernetes, Docker, cloud infrastructure, and GPU-based workloads.
  • Strong understanding of distributed systems and large-scale ML infrastructure design.
  • Experience with ML workflow orchestration systems such as Ray, Kubeflow, Argo, Airflow, Flyte, or Metaflow.
  • Experience deploying and managing production inference systems using platforms like Triton, vLLM, TensorRT-LLM, Ray Serve, KServe, Seldon, BentoML, TorchServe, or custom services.
  • Strong understanding of inference optimization techniques such as batching, quantization, CUDA graphs, and memory-aware scheduling.
  • Experience with model registries, experiment tracking, CI/CD for ML, canary deployments, shadow traffic, rollback strategies, and production monitoring.
  • Strong cloud experience across AWS, GCP, Azure, or GPU-focused providers like CoreWeave, Lambda Labs, or RunPod.
  • Ability to debug performance bottlenecks across distributed systems, containers, networking, GPU memory, and storage layers.

Strong ownership mindset with the ability to define architecture, set platform standards, and drive execution across teams.

Nice to Have

  • Experience with multimodal, vision, or generative AI systems.
  • Experience with large-scale GPU clusters e.g. A100/H100, NCCL, and high-throughput data pipelines.
  • Experience designing evaluation and monitoring systems for generative AI workloads.
  • Familiarity with ML security, privacy, and data governance practices.
  • Experience building internal developer platforms for research teams.


Success Looks Like

  • Within 6 months, you will:
  • Create reliable research-to-production pathways for SPREEAI’s core AI models.
  • Reduce manual model deployment friction through standardized pipelines and tooling.
  • Improve GPU utilization and reduce training and inference costs.
  • Establish robust observability and evaluation gates for production model releases.
  • Accelerate the delivery of new AI capabilities into partner-facing experiences.


Why This Role Matters

This is not a traditional DevOps role. This is the infrastructure backbone that enables SPREEAI to turn frontier AI research into reliable, scalable, production-grade systems. You will define the systems powering real-time AI experiences where latency, cost, and model quality directly impact end-user experience.


Why Join SPREEAI?

  • Build the Core AI Infrastructure, Not Just Features: You will define how multimodal AI systems are reliably deployed, monitored, and scaled—directly shaping the performance, cost efficiency, and reliability of real-world AI products.
  • Own Systems End-to-End: You will own critical infrastructure decisions across deployment, observability, and resource management, with direct impact on production systems serving real partner traffic.
  • Work on Hard, High-Leverage Problems: From GPU efficiency to large-scale deployment systems, you will tackle challenges that sit at the frontier of real-time AI infrastructure.
  • High Velocity, Low Bureaucracy & Direct Impact: We operate with tight feedback loops between research, platform, and product, enabling rapid iteration and meaningful impact without organizational friction.

Similar Jobs

5 Minutes Ago
In-Office
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Artificial Intelligence • Hardware • Productivity • Robotics • Software • Automation • Manufacturing
Own the full employee experience as the sole HR Generalist, including onboarding, offboarding, employee relations, engagement, payroll, benefits, leave administration, compliance, audits, HRIS management, reporting, and process automation. The role requires building scalable HR systems from scratch, supporting managers through sensitive personnel decisions, and maintaining compliance across multiple states. This position is onsite four days per week at the Carson headquarters.
Top Skills: AirtableAshbyClaudeClickupEthenaHrisNotionScribeSlackUkgWex
8 Minutes Ago
In-Office or Remote
7 Locations
153K-270K Annually
Senior level
153K-270K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Build privacy-focused services and infrastructure that support consumer privacy rights across Block. Collaborate with engineering, security, platform, and legal teams to embed privacy solutions, define future privacy technology, mentor engineers, review implementations, and maintain monitoring and system stability. The role involves developing scalable, testable software and creating architectures and processes addressing data governance, data flows, access control, security, and privacy risks.
Top Skills: AWSJavaKotlinLinuxMySQLTerraform
8 Minutes Ago
In-Office or Remote
7 Locations
173K-312K Annually
Senior level
173K-312K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Leads Sales Development across North America and the UK, owning top-of-funnel pipeline performance, team integration, reporting, targeting, conversion, and revenue contribution. Builds scalable inbound and outbound motions, expands upmarket coverage, develops managers and representatives, and partners with Marketing, Sales, Finance, Risk, Compliance, Product, and Operations. Uses analytics, automation, and AI to improve prospecting and productivity while establishing rigorous planning, coaching, qualification, and performance standards.
Top Skills: AIAnalyticsAutomationReporting Tools

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account