Hands-on Senior AI Engineer to design and ship production Generative AI systems: build RAG pipelines, embeddings, vector search, multi-agent orchestration, backend APIs (Python/Flask/FastAPI), deploy cloud-native workloads, implement LLMOps practices, and mentor through code reviews while owning end-to-end delivery.
We're looking for a hands-on AI Engineer who combines strong backend engineering fundamentals with hands-on experience building production Generative AI systems. You'll design and ship RAG pipelines, integrate LLMs into real products, and build the backend services that support them, writing code daily, not just architecting on paper.
This is a purely technical IC role, not a managerial one. You’ll lead by example, mentor through code reviews, and own end-to-end technical delivery.
Key Responsibilities
- Design and build RAG systems, embeddings, vector search, chunking, and evaluation pipelines.
- Build and maintain multi-agent orchestration workflows (LangGraph, AutoGen, CrewAI, or similar).
- Develop backend services and APIs (Python — Flask/FastAPI) that expose AI workflows to production systems.
- Deploy and scale AI workloads in cloud-native environments, using serverless or containerized patterns.
- Implement LLMOps practices: prompt versioning, cost tracking, monitoring, and evaluation.
- Write clean, tested code, and use AI-assisted tools (Copilot, Cursor, Claude Code) to move faster without cutting corners.
- Work with data and platform engineers to ship GenAI features quickly, from prototype to production.
Skills, Knowledge and Expertise
Must-Have Skills
- 5+ years of backend experience, with strong Python coding skills.
- Proven experience shipping RAG systems (vector DBs, embeddings, chunking).
- Familiarity with orchestration frameworks (LangGraph, LangChain, AutoGen, or similar).
- Experience with APIs, microservices, and cloud-native development (AWS preferred).
- Familiarity with distributed systems concepts (async, message queues, caching).
- Experience with unstructured data (PDFs, tables, images).
- Builder mindset: thrives on writing, debugging, and improving production code.
- Collaborative, humble, and open to feedback.
- Strong communicator who explains design decisions clearly.
Influences through contribution, not hierarchy.
About
Emumba is a global engineering and consulting company with strengths in software development and an established AWS cloud practice focused on Data and GenAI. For 15 years, our teams across the US, the UAE, and Pakistan have earned trust through quality delivery and ownership of work. We look for people who value the culture they work in as much as the craft they bring to it.
Emumba Santa Clara, California, USA Office
Santa Clara, CA, United States
Similar Jobs
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Design, build, and deploy enterprise-scale generative AI and LLM-powered applications and agentic workflows. Implement RAG pipelines, document ingestion, embeddings, semantic and hybrid search, and integrate vector databases with PostgreSQL. Build responsive frontends (React/Next.js) and backend services (Python/Node.js), define end-to-end architecture, lead technical decisions and reviews, mentor engineers, and ensure safe, scalable AI solutions in collaboration with product and business partners.
Top Skills:
Agentic WorkflowsAi Orchestration FrameworksCrewaiEmbeddingsGenerative AiGraphragHybrid SearchLangchainLanggraphLlmsMicroservicesNext.JsNode.jsPostgresPythonReactRest ApisRetrieval-Augmented Generation (Rag)Semantic SearchVector Databases
AdTech • Artificial Intelligence • Big Data • Machine Learning • Marketing Tech • Mobile • Software
Design, prototype, and productionize generative AI features and agentic workflows. Build Python services, TypeScript/React UIs, data pipelines, evaluation and instrumentation, and safeguards. Partner with product and stakeholders to measure impact, ensure reliability, and mentor engineers.
Top Skills:
Generative AiLarge Language ModelsPythonReactSQLTypescript
Edtech • Healthtech • HR Tech • Information Technology • Professional Services • Software • Telehealth
Design, build, test, deploy, and maintain cloud-based web applications and microservices. Ensure performance, reliability, and code quality while collaborating on new features, debugging, and optimizing systems. Use AI development tools daily, contribute to infrastructure as code, and promote maintainable architecture and documentation.
Top Skills:
.Net.Net CoreAngularAWSAzureAzure DevopsC#Ci/CdClaudeCursorDockerGitGithub CopilotHTML/CSSJavaScriptKubernetesMicroservicesNgrxOauthRedisSQLTerraformTypescriptWeb ApiXunit
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



