Our mission is to free humanity from meaningless work. We are building the System of Record for Enterprise Compliance, turning hours of manual, high-stakes marketing workflows into minutes of automated precision.
We are a small, high-density team of engineers and operators. Having proven our value with global brands, we are now at the inflection point where our technical architecture meets massive scale. This is a "rocket ship" moment: we are moving beyond simple automation into a world of vision-first, agentic workflows that solve the problems generic frontier models cannot.
The Role: Architect of the Agentic Brain
As Senior Applied AI Engineer, you’ll own the technical moat that makes us stand alone: the safety dataset and deterministic orchestration that prevents enterprise brands from ever trusting generic AI with compliance. You’re not building ‘better automation’—you’re building the category-defining infrastructure that gives us a 3-year lead in a market where second place doesn’t exist.
You will lead the technical evolution of our agent-driven architecture: defining the role of each agent, designing multi-step reasoning flows, integrating tools, memory, and retrieval systems, and optimizing for accuracy, determinism, and trust. Your work will directly determine whether enterprise customers can rely on Puntt for legal and brand compliance at scale.
This is a hands-on, in-office role in San Francisco, working closely with a small, senior team to build systems where correctness matters more than demos.
What You’ll Own
Agentic Reasoning & Orchestration
Design and evolve multi-agent LLM systems that decompose complex review tasks into reliable, auditable steps.
Define agent responsibilities, hand-offs, and termination conditions to minimize reasoning drift and maximize consistency.
Context, Retrieval & Memory Systems
Architect retrieval pipelines using RAG, structured memory, and emerging approaches like graph-based retrieval to provide agents with the right context at the right time.
Balance recall, precision, and latency across large knowledge bases (brand guidelines, regulations, historical decisions).
Stateful, Asynchronous Workflows
Own long-running, fault-tolerant workflows using Temporal (or similar), ensuring retries, versioning, and determinism across non-deterministic model calls.
Treat agent orchestration as a distributed systems problem: managing state, failures, and observability.
Evaluation, Safety & Reliability
Build evaluation frameworks that go beyond “it looks right,” using statistical metrics, gold labels, and automated regression testing to prove system reliability.
Prioritize correctness and trust, especially in high-risk legal and compliance scenarios.
Asset Understanding Pipeline
Collaborate on image and document preprocessing (OCR, layout analysis, VLMs) to ensure downstream agents receive structured, machine-readable context.
Focus on practical understanding, not computer vision research.
End-to-End Ownership
Move fluidly between Python-based LLM services, retrieval pipelines, and AWS infrastructure to ship reliable systems end-to-end.
Who You Are: The Hybrid Systems Builder
You are someone who enjoys building real systems with LLMs, not just experimenting with them.
Strong Engineering Foundation
You have 5–7+ years of experience building production systems and understand core CS concepts—data structures, concurrency, failure modes, and tradeoffs.
Experienced with LLM-Driven Systems
You’ve spent 1–2+ years building with large language models in real applications: tool use, function calling, structured outputs, and multi-step reasoning.
Agentic & Retrieval-First Thinker
You’re comfortable designing systems that combine LLMs with RAG, memory, graph-based context, and external tools rather than relying on a single prompt.
Systems-Oriented
You see multi-agent orchestration as a distributed systems challenge—latency, retries, observability, and consistency all matter.
Comfortable with Ambiguity
You thrive in an early-stage environment where problems are underspecified and the best solution doesn’t exist yet.
Technical Requirements
Must-have:
5–7+ years of professional engineering experience, with a strong record of shipping production systems
1–2+ years building with LLMs in real applications (not just experimentation)
Expert Python experience
Hands-on experience designing RAG systems, vector search, embeddings, and structured retrieval
*Preferred:
Experience with LLM orchestration frameworks (e.g., LangGraph, CrewAI, or custom orchestration layers)
Experience with stateful workflow orchestration (Temporal a plus)
Experience operating AI systems on AWS (Lambda, S3, Bedrock, SageMaker, etc.)
Strong TypeScript experience*
Bonus: experience with OCR, document parsing, or VLMs
Why Join Puntt
Small Team, Real Ownership
You will be a foundational technical leader shaping how the system works, not just implementing tickets.
High-Impact, High-Trust Domain
You’re building AI systems where correctness matters—and where most “generic AI” solutions fail.
Speed Without Chaos
We ship quickly, but we care deeply about system design, evaluation, and long-term reliability.
Similar Jobs
What you need to know about the San Francisco Tech Scene
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



