Quantiphi Logo

Quantiphi

Technical Architect - ML

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Architect and implement enterprise-grade MLOps/LLMOps platforms on AWS and Kubernetes. Design ML/LLM pipelines for training, validation, deployment, monitoring, governance, CI/CD, and observability. Serve as technical authority, conduct architecture reviews, ensure security and compliance, and mentor engineering teams across projects.
The summary above was generated by AI

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Must have skills & Qualifications:

  • 8+ years working in ML/AI engineering or MLOps roles with strong architecture exposure.

  • Strong expertise in AWS cloud-native ML stack, including: SageMaker(primary), EKS, Lambda, API Gateway, CI/CD (CodeBuild/CodePipeline or equivalent)

  • Hands-on experience with at least one major MLOps toolset and awareness of alternatives: MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, Seldon.

  • Deep understanding of model lifecycle management (feature engineering->training → registry → deployment → monitoring).

  • Experience implementing or supporting LLMOps pipelines, including: prompt versioning, evaluation metrics, automation frameworks

  • Deep understanding of ML lifecycle: data ingestion, feature engineering, training, evaluation, model packaging, CI/CD, drift detection, monitoring, and governance.

  • Strong experience with AWS SageMaker (Pipelines, Feature Store, Model Registry, Model Monitor).

  • Experience implementing ML CI/CD pipelines including automated training, testing, validation, model promotion, and endpoint deployment.

  • Experience working on Infrastructure as Code (IaC) tools and CI/CD pipelines

  • Experience with Kubernetes based development

  • Experience with feature engineering pipelines and Feature Store management.

  • Understanding of lineage tracking: training data snapshot, feature versions, code versioning, metadata tracking, reproducibility.

  • Hands-on experience with AWS Bedrock and Agentcore service

  • Experience with CloudWatch, SageMaker Model Monitor, Prometheus/Grafana.

  • Strong foundation in Python and cloud-native development patterns.

  • Solid understanding of security best practices, IAM, secrets management, and artifact governance.

Good to have skills:

  • Experience with vector databases, RAG pipelines, or multi-agent AI systems.

  • Exposure to DevOps and infrastructure-as-code (Terraform, Helm, CDK).

  • Hands-on understanding of model drift detection, A/B testing, canary rollouts, and blue-green deployments.

  • Familiarity with Observability stacks (Prometheus, Grafana, CloudWatch, OpenTelemetry).

  • SQL and data transformation experience using Snowflake, Databricks, Spark.

  • Ability to translate business goals into scalable AI/ML platform designs.

  • Strong communication and cross-team collaboration skills.

  • Ability to guide engineering teams through technical uncertainty and design choices.

Key Responsibilities:

  • Architect and implement the MLOps strategy for the programme, ensuring alignment with the project proposal and delivery roadmap.

  • Design and own enterprise-grade ML/LLM pipelines covering model training, validation, deployment, versioning, monitoring, and CI/CD automation.

  • Build container-oriented ML platforms (EKS-first) while evaluating alternative orchestration tools with similar capabilities (Kubeflow, SageMaker, MLflow, Airflow, etc.).

  • Implement hybrid MLOps + LLMOps workflows, including prompt/version governance, evaluation frameworks, and monitoring for LLM-based systems.

  • Serve as a technical authority across multiple internal and customer projects, contributing architectural patterns, best practices, and reusable frameworks.

  • Enable observability, monitoring, drift detection, lineage tracking, and auditability across ML/LLM systems.

  • Define and implement standards for model deployment, monitoring, governance, and automation to ensure production-grade reliability and scalability.

  • Collaborate with cross-functional teams — data engineering, platform, DevOps, and client stakeholders — to deliver production-ready ML solutions.

  • Ensure all solutions adhere to security, governance, and compliance expectations, particularly around handling cloud services, Kubernetes workloads, and MLOps tools.

  • Conduct architecture reviews, troubleshoot complex ML system issues, and guide teams through implementation across cloud-native ML platforms.

  • Mentor engineers and provide guidance on modern MLOps tools, platform capabilities, and best practices.

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Quantiphi San Jose, California, USA Office

2460 N 1st Street, San Jose, San Jose, United States, 95131

Similar Jobs

Yesterday
Remote
USA
Senior level
Senior level
Artificial Intelligence • Big Data • Machine Learning
Design and deliver enterprise-grade Generative AI solutions on AWS using Bedrock and AgentCore. Architect LLM-based applications, RAG pipelines, agentic workflows, vector DB integrations, and APIs. Optimize models for performance, cost, and quality, ensure security and governance, lead teams, and troubleshoot production GenAI systems.
Top Skills: Amazon AgentcoreAmazon Api GatewayAmazon S3Apache AirflowAws BedrockAws LambdaAws SagemakerAws Step FunctionsCi/CdEmbeddingsIacKubeflowLangchainLlm Apis (ClaudeModel Context Protocol (Mcp)Nova)Prompt EngineeringRetrieval-Augmented Generation (Rag)Sagemaker PipelinesStrand AgentsVector Databases
18 Days Ago
In-Office or Remote
196K-257K Annually
Expert/Leader
196K-257K Annually
Expert/Leader
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Database • Analytics
Serve as a technical leader for legacy-to-AI cloud migrations and enterprise generative AI architecture. Create reusable solution patterns and reference architectures, advise C-suite stakeholders, architect LLM robustness and guardrails, and manage high-complexity data migrations into Snowflake AI workloads.
Top Skills: ExadataGenerative AiLarge Language Models (Llms)Llm FrameworksNetezzaPrompt Safety / GuardrailsSnowflakeSnowflake CortexTeradataVector Data Types
10 Hours Ago
In-Office or Remote
Junior
Junior
Angel or VC Firm • Professional Services • Consulting • Financial Services
Join a boutique investment banking team to build and own complex financial models, prepare client-ready presentations, drive M&A and capital-raising processes, manage diligence and coordination with advisors, and interact directly with senior clients. Requires strong analytical, quantitative, and client communication skills and elite execution on lean deal teams.
Top Skills: Claude PluginExcelGaapIndex/MatchMacrosPivot TablesPower QueryPowerPointVlookupWordXlookup

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account