Avantos.ai Logo

Avantos.ai

Senior DevOps Engineer

Sorry, this job was removed at 03:11 p.m. (PST) on Wednesday, Jun 03, 2026
Remote
Hiring Remotely in United States
Remote
Hiring Remotely in United States

Similar Jobs

4 Days Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Fintech • Software • Financial Services
Build and operate scalable, reliable infrastructure using Terraform, Kubernetes, AWS, and CI/CD automation. Responsibilities include managing EKS platforms, improving deployment pipelines, securing networking and IAM, enhancing observability, optimizing cloud costs, implementing disaster recovery, reducing operational toil, and modernizing legacy infrastructure. The role also leads incident response, reliability initiatives, platform adoption, and cross-team infrastructure projects.
Top Skills: ArgocdAWSBashDatadogEksGithub ActionsIamJavaScriptKafkaKubernetesLambdaMskPostgresPythonRdsRedisS3TerraformTypescriptVpc
5 Days Ago
Remote or Hybrid
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Fintech • Mobile • Social Impact • Software • Financial Services
Deploys, monitors, and maintains production infrastructure and services. Responsibilities include developing infrastructure and network best practices, supporting on-call operations, operating SDLC and CI/CD platforms, improving system performance, providing operational support for distributed applications, and analyzing uptime and production performance. The role requires AWS administration, DevOps or SRE experience, version control, scripting, Terraform, containerized AWS services, Linux, networking, and security-sensitive environment experience.
Top Skills: AWSAws CodepipelineBashCi/CdClaudeDevOpsDnsDockerGitGitlabHttp/HttpsJenkinsKubernetesLinuxMlopsPostgresPythonScpSftpSreSshSysopsTerraform
Yesterday
Remote
United States
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Own enterprise deployments of Qodo’s Kubernetes-based platform in customer-managed AWS, GCP, and Azure environments. Install, configure, upgrade, troubleshoot, and optimize infrastructure, writing Python or Bash diagnostics and codifying repeatable processes with Terraform and Helm. Partner directly with customer engineering teams, resolve complex networking and access issues, maintain PostgreSQL and Redis infrastructure, and improve deployment reliability through automation and documentation.
Top Skills: AWSAzureBashDockerGCPHelmIamKubernetesLinuxPostgresPythonRedisTerraform

Company overview

Avantos is building the industry’s first AI-native operating system for financial services, redefining how firms onboard clients, deliver advice, and manage core servicing workflows. Our platform unifies fragmented data, automates complex processes, and embeds intelligent decision-making across every step of the client lifecycle.

We partner with leading financial institutions and are scaling rapidly. We’re an execution-driven, design-obsessed, product-led team composed of founders and leaders from Wharton, MIT, top design programs, and prior unicorn SaaS companies. We move fast, solve deep industry problems, and build technology that puts users back in control of their workflows.

If you love client impactproduct designcomplex problem solving, and bringing AI-enabled change to real-world businesses, Avantos is where you will thrive.

Job summary

We're seeking a Senior DevOps Engineer / Site Reliability Engineer to own and evolve our infrastructure, reliability, and deployment practices. You'll be responsible for building the foundational platform that enables our engineering teams to ship quickly and reliably while maintaining the security and compliance standards required in financial services.

  • Design, implement, and maintain our AWS cloud infrastructure using infrastructure-as-code principles with Terraform
  • Build and optimize CI/CD pipelines to enable rapid, safe deployments across multiple environments
  • Own observability strategy—implement comprehensive monitoring, logging, and alerting systems using Datadog and other tooling
  • Architect and manage containerized workloads on ECS Fargate and evaluate migration paths to Kubernetes
  • Establish and enforce security best practices, working closely with compliance teams on financial services requirements
  • Design and implement disaster recovery, backup, and business continuity strategies
  • Optimize system performance, cost efficiency, and resource utilization across AWS services
  • Collaborate with engineering teams to improve service reliability, reduce toil, and establish SLOs/SLIs
  • Participate in incident response and conduct thorough post-mortems to drive continuous improvement
  • Mentor engineers on DevOps practices, cloud architecture patterns, and operational excellence

Your skills will include

  • 8+ years of experience in DevOps, SRE, or infrastructure engineering roles
  • Expert-level proficiency with AWS services including ECS Fargate, ALB, Cognito, S3, SQS, and related services
  • Deep hands-on experience with Terraform for managing complex, multi-account AWS environments
  • Strong scripting and automation skills in Python and/or Bash
  • Proven experience designing and implementing CI/CD pipelines (GitHub Actions, ArgoCD, or similar)
  • Solid understanding of containerization technologies (Docker) and orchestration platforms (Kubernetes/ECS)
  • Experience with observability and monitoring tools (Datadog, CloudWatch, or equivalent)
  • Deep knowledge of networking, security, and AWS best practices
  • Strong problem-solving abilities and experience troubleshooting complex distributed systems
  • Excellent communication skills and ability to work cross-functionally with engineering teams

Nice to haves

  • Experience in financial services or highly regulated industries
  • Familiarity with event-driven architectures and message queue systems (Kafka, SQS)
  • Experience with PostgreSQL performance tuning and RDS management
  • Knowledge of microservices architecture patterns and service mesh technologies
  • Experience with security tooling, vulnerability scanning, and compliance frameworks
  • Familiarity with our application stack (Golang, Next.js, PostgreSQL)
  • Experience managing AI/ML infrastructure and AWS Bedrock

What we offer:

  • Competitive compensation + meaningful equity
  • Opportunity to build production infrastructure from the ground up for a rapidly scaling AI platform
  • A culture optimized for engineering excellence, focus, deep work, and ownership—not ticket factories
  • Remote work flexibility

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account