Nace.AI Logo

Nace.AI

Senior MLOps Engineer

Reposted 3 Hours Ago
In-Office
Palo Alto, CA, USA
Senior level
In-Office
Palo Alto, CA, USA
Senior level
Own and operate end-to-end ML infrastructure for training, serving, and evaluation of LLM/SLM models. Build scalable low-latency inference (vLLM, batching, autoscaling), multi-GPU training/serving clusters, observability and monitoring, inference-time optimizations (quantization, distillation), reproducibility, versioning, and enterprise-grade auditability. Set MLOps best practices and standards.
The summary above was generated by AI

Palo Alto, CA | Full-Time | On-site

About Nace AI:

Nace AI is an enterprise AI product and research company in Palo Alto (backed by General Catalyst, Walden Catalyst, and Intel). We build long-running AI agents powered by our own specialized SLMs — we started with financial audit and accounting workflows and are expanding from there. Real enterprise deployments, not demos.

Role Overview:

As a Senior MLOps Engineer, you will own the infrastructure that takes Nace.AI's models from research to reliable, production-grade systems. Our infrastructure generates task-specific Small Language Models (SLMs) in real time — which means our training, serving, and evaluation infrastructure isn't an afterthought; it is the product. You will design and operate the pipelines, orchestration, and serving layers that allow us to train, deploy, monitor, and continuously improve many specialized models at once, with the reliability that high-stakes audit, compliance, and finance workflows demand. This role sits at the intersection of ML engineering, LLM inference infrastructure, and platform reliability, and requires both strong systems instincts and hands-on execution.

Key Responsibilities:

  • Design, build, and operate end-to-end ML infrastructure: training orchestration, experiment tracking, model registries, CI/CD for models, and automated evaluation pipelines.

  • Own LLM/SLM serving infrastructure — scale low-latency, high-throughput inference using frameworks like vLLM, including batching, caching, and autoscaling strategies.

  • Build and manage multi-GPU training and inference clusters (scheduling, utilization, cost optimization) across cloud and on-prem environments.

  • Implement observability for models in production: latency, throughput, drift, regression, and quality monitoring with actionable alerting.

  • Apply inference-time optimizations — quantization (AWQ, GPTQ, FP8/GGUF), distillation support, KV-cache management, and deployment tuning — in partnership with our ML and Research Engineers.

  • Harden our stack for enterprise deployment: reproducibility, versioning, access controls, and audit-ready traceability of model behavior.

  • Set MLOps best practices and tooling standards as an early, senior member of the infrastructure team.

Qualifications:

  • 5+ years of experience in MLOps, ML infrastructure, or platform engineering, with substantial production ownership.

  • Proven experience deploying and scaling LLM, inference infrastructure in production, including model serving frameworks such as TRT, vLLM, SGLang or TGI.

  • Strong proficiency with Kubernetes, containerization (Docker), and infrastructure-as-code (Terraform or similar).

  • Hands-on experience with GPU cluster management and distributed training/serving environments.

  • Proficient in Python with a strong track record of building substantial, maintainable systems.

  • Experience with ML pipeline and orchestration tooling (e.g., Airflow, Kubeflow, Ray, MLflow, Weights & Biases).

  • Solid foundation in computer science fundamentals and cloud architecture (AWS, GCP, or Azure).

  • BS degree in CS or related technical field.

  • Self-starter comfortable working in a fast-paced, dynamic environment.

Preferred Qualifications:

  • MS in CS or related technical field.

  • Experience operating multi-node GPU training infrastructure.

  • Hands-on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF) and other inference-time optimizations.

  • Familiarity with data processing stacks such as Spark and Airflow.

  • Experience supporting fine-tuning workflows for LLMs/VLMs (instruction tuning, RLHF/DPO pipelines).

  • Experience in regulated or enterprise environments where reliability, security, and auditability are first-class requirements.

  • Contributor to open-source ML infrastructure projects.

Why Nace AI?

  • Pedigree: Work with a team from top-tier institutions and companies, backed by the best VCs in the world.

  • Impact: You are joining early enough to shape the infrastructure foundations of a company aiming to be the "OS" for professional knowledge.

  • Competitive Package: Silicon Valley-standard salary, significant equity, and premium benefits.

Similar Jobs

5 Days Ago
In-Office
Santa Clara, CA, USA
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and operate end-to-end cloud data and ML pipelines that ingest, validate, process, label, and evaluate multimodal sensor data for autonomous driving. Own architecture, reliability, observability, and operational metrics; collaborate with perception, ML, labeling, and product teams; provide technical leadership, code contributions, and mentorship to deliver AV-scale systems.
Top Skills: 3D GeometryC++CameraCi/CdCloud InfrastructureComputer VisionData PlatformsDeep LearningDistributed SystemsGpu-Accelerated ComputingLidarMlopsPerception PipelinesPythonRadarWorkflow Orchestration
6 Days Ago
In-Office
Sunnyvale, CA, USA
Senior level
Senior level
Healthtech • Robotics
Design, build, and maintain production MLOps infrastructure: Kubernetes clusters, ML orchestration, GPU validation, storage and CI/CD integration, migrations, runbooks, on-call incident response, and security/compliance collaboration to support reproducible ML workflows at scale.
Top Skills: A6000AnsibleArgocdB200BashCniCudaGitlab CiHelmKubeflowKubernetesL40SLinuxMetaflowMigMinioMlflowNetappNumaNvidia DriversNvlinkPythonS3TerraformV100
16 Days Ago
In-Office
San Jose, CA, USA
149K-216K Annually
Senior level
149K-216K Annually
Senior level
Artificial Intelligence • Internet of Things • Machine Learning
Design, build, and operate scalable ML pipelines and infrastructure across cloud and on‑prem HPC. Implement MLOps tooling (tracking, registries, feature stores), CI/CD/CT, containerized GPU orchestration, model optimization (LLMs, GNNs, RL), data/versioning pipelines, monitoring, and mentor engineers to productionize ML for EDA and simulation workloads.
Top Skills: AirflowArizeAutogenAws SagemakerAzure MlBashCadenceCloudFormationDeepspeedDockerDvcElk StackEvidently AiFeastFsdpGcp Vertex AiGoGrafanaHugging FaceJaxKubeflowKubernetesLangchainLsfMlflowPrometheusPythonPyTorchScikit-LearnSiemens EdaSlurmSQLSynopsysTensorFlowTerraformWeights & BiasesXgboost

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account