Schemata Logo

Schemata

LLM Platform Engineer

Posted 10 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Mid level
In-Office
San Francisco, CA, USA
Mid level
Build and operate multimodal LLM platform capabilities, including knowledge ingestion, embeddings, retrieval-augmented generation, evaluation, feedback loops, deployment, monitoring, and latency optimization. The role spans cloud, on-premises, and air-gapped environments, integrating AI knowledge systems with 3D scene representations and real-time training applications. Responsibilities include productionizing AI updates, improving reliability, and collaborating with spatial computing, mobile, and product engineering teams.
The summary above was generated by AI

About Schemata

At Schemata, we are transforming the $400 B virtual‑training and simulation market by fusing 3D computer vision, neural rendering and large multimodal models inside highly regulated industries. Our platform delivers photorealistic, intelligent 3D experiences, demanding robust spatial reasoning, high‑performance data pipelines and seamless integration between traditional graphics and AI‑driven perception.

About the Role

We are seeking a highly skilled LLM Platform Engineer to join our team full‑time. You will play a foundational role in designing, building and optimizing the AI systems that turn heterogeneous knowledge — technical documentation, video, imagery, voice, and 3D scene data — into grounded, real‑time guidance for next‑generation training and maintenance applications.

This is a high‑impact, cross‑functional role: you will work end‑to‑end from evaluating emerging methods to production inference and performance optimization, ensuring our platform retrieves the right knowledge, reasons over it reliably, and responds in real time across diverse deployment environments.

Core Responsibilities
  • Design and build multimodal knowledge base ingestion: pipelines that convert documents, video, images, audio, and structured data into a unified, queryable semantic layer with rich embeddings and metadata, including bringing in external sources and messy real‑world formats as they arise.

  • Improve our retrieval‑augmented generation (RAG) systems: chunking and indexing strategies, hybrid retrieval, reranking, grounding controls, and context assembly for multimodal queries.

  • Build the knowledge base lifecycle: handle versioned documentation updates, incorporate SME corrections and annotations, and design validation and freshness mechanisms so the knowledge base improves continuously.

  • Scale the feedback loop: capture user signals (ratings, corrections, query patterns), route them into measurable improvements to retrieval and content, and build a durable picture of how the AI is performing in the field.

  • Own the deployment and delivery lifecycle of AI system updates across cloud, on‑premises, and air‑gapped environments: packaging, release management, configuration, and monitoring.

  • Strengthen the reliability and latency of AI features: profiling inference paths, caching, fallback behavior, and graceful degradation in production.

  • Collaborate with spatial computing, mobile, and product engineers to connect the knowledge base with 3D scene representations and ship mission‑critical features.

Essential Skills & Experience
  • 4+ years of software engineering experience, with at least 1–2 years building production LLM systems (RAG pipelines, agents, or LLM‑powered products with real users).

  • Hands‑on experience designing knowledge ingestion and retrieval systems: embedding models, vector and hybrid search, chunking strategies, and semantic data modeling across more than one modality.

  • Strong Python and backend engineering fundamentals: APIs, data pipelines, async systems, and working in a cloud environment (AWS preferred).

  • A metrics‑driven approach: experience building evals or benchmarks for LLM systems and using them to drive iteration, not just report scores.

  • Demonstrated ability to take ambiguous problems from prototype to reliable, maintainable production services.

  • Experience deploying and operating production systems: CI/CD, release management, and monitoring.

Nice to Have
  • Experience fine‑tuning open‑weight models (LoRA/PEFT, distillation) or deploying quantized models for offline, edge, or on‑premises environments.

  • Research background or publications in retrieval, multimodal learning, or agentic systems, or a track record of translating recent research into shipped capabilities.

  • Voice pipeline experience: streaming STT/TTS, latency optimization for real‑time conversational systems.

  • Familiarity with 3D or spatial data (scene graphs, point clouds, 3DGS) and grounding language models in spatial representations.

  • MLOps and inference infrastructure experience: GPU serving, model versioning, cost/latency optimization at scale.

  • Defense, aerospace, energy or other regulated‑industry experience; active or ability to obtain U.S. security clearance.

Why Join Us?

  • Competitive salary that reflects your experience and track record

  • Meaningful equity stake in a high-growth, venture-backed defense tech startup, so you share in the upside you help create

  • Comprehensive health coverage: medical, dental, and vision insurance

  • 401(k) plan

  • Paid parental leave

  • High visibility and real impact: Collaborate with world-class engineers and researchers in a high-ownership environment.

Similar Jobs

One Month Ago
In-Office
San Francisco, CA, USA
140K-220K Annually
Entry level
140K-220K Annually
Entry level
Artificial Intelligence • Gaming • Software
Own the infrastructure powering Pax Historia’s AI requests across 37+ models and multiple providers. Responsibilities include improving provider reliability, structured outputs, caching, routing, monitoring, cost efficiency, latency, and scalability for billions of monthly tokens. This founding engineer role requires strong fundamentals, attention to detail, data-driven decision-making, and rapid learning. The position is fully in-person in San Francisco and does not involve model training or self-hosted inference.
Top Skills: Ai Provider ApisCaching SystemsLlmMonitoring SystemsRouting SystemsStructured Outputs
One Month Ago
In-Office
San Francisco, CA, USA
176K-220K Annually
Senior level
176K-220K Annually
Senior level
Edtech • Enterprise Web • HR Tech • Software
Build and operate shared ML/AI infrastructure: data pipelines, feature stores, training, model serving and LLM platform work. Scale inference and training (GPU, batching, autoscaling), enable evaluation and post-training workflows, and partner with AI, data science, and product teams to productionize models and improve platform reliability and developer experience.
Top Skills: AirflowApache BeamAutoscalingAWSBatchingBigQueryCi/CdDataflowDockerEmbeddingsFeature StoresGCPGoGpu InferenceKubernetesLlmsModel ServingPythonSparkStreaming PipelinesTerraformTypescript
One Month Ago
In-Office
San Francisco, CA, USA
200K-345K Annually
Mid level
200K-345K Annually
Mid level
eCommerce • Mobile • Retail
The LLM Platform Engineer will design scalable AI/ML infrastructures, develop frameworks for model evaluation, and bridge research with production solutions while collaborating with machine learning scientists.
Top Skills: Apache KafkaAws Ec2Aws EcsAws EksAws KinesisAws LambdaAws S3Aws SagemakerDatadogDynamoDBElasticsearchFlinkGrafanaPostgresPythonRedis

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account