Collective Logo

Collective

Senior Cloud Infrastructure Engineer

Posted 5 Days Ago
Be an Early Applicant
Hybrid
San Francisco, CA, USA
200K-230K Annually
Senior level
Hybrid
San Francisco, CA, USA
200K-230K Annually
Senior level
Own and evolve Collective’s multi-cloud infrastructure across AWS and GCP using Terraform. Lead CI/CD, observability, identity, security, disaster recovery, automation, incident response, and platform reliability. Partner with product engineering and AI Platform teams to deliver scalable infrastructure, reduce operational toil, improve system performance, and mentor engineers. Participate in on-call rotations and drive post-incident improvements.
The summary above was generated by AI

About Collective:

Collective is on a mission to redefine the way businesses-of-one work. Our technology and team of trusted advisors help members achieve financial independence by taking care of everything from business incorporation to accounting, bookkeeping, tax services, and access to a thriving community, all in one integrated platform. We believe in empowering self-employed people to enjoy the same tax savings that big companies get, so they can focus on their passion, not paperwork.

Featured in Forbes, Business Insider, Yahoo, Bloomberg, Financial Times, TechCrunch, and more. We are backed by General Catalyst, Sound Ventures, QED Investors, Google’s Gradient Ventures, Expa, and other investors who have financed iconic companies like YouTube, Substack, Twitch, Box, Hims, Instacart, and Lyft.

About the role:

You'll join Collective's Infrastructure team, a small group that owns the platform every engineer here builds on: AWS and GCP foundations, CI/CD, IaC, observability, secrets management, and identity. We partner with product engineering to unblock shipping, with the security team to keep the platform safe, and with the AI Platform team on the infra behind Collective's growing AI investments. As a Senior Cloud Infrastructure Engineer on this team, you'll pave new paths, test recovery plans, and manage quiet, stable systems. You'll be the person justifying design choices with production experience and the one others turn to when a system needs rethinking.

What you'll do: 

  • Use Infrastructure as Code (IaC) with Terraform to provision, deploy, and manage cloud resources on AWS and GCP as well as other SaaS vendors.

  • Embed security best practices into the infrastructure by enforcing zero-trust architecture principles like least privilege and identity-based access to protect systems and data.

  • Build scalable, reliable, and cost-effective systems that hold up as Collective grows.

  • Develop and test disaster recovery plans.

  • Own the CI/CD system engineering teams ship on. Set standards, drive reliability and speed improvements, and mentor teams on best practices.

  • Reduce operational toil across the platform through automation, leveraging AI tooling where it accelerates safe, high-quality work.

  • Work closely with product engineering teams to understand application needs and translate them into scalable infrastructure solutions.

  • Own the observability stack (monitoring, logging, alerting) and use it to proactively identify and remediate performance bottlenecks.

  • Participate in the on-call rotation to respond to outages, recover systems, own incident response and post-mortem.

  • Stay current with emerging technologies and best practices in Cloud Infrastructure, DevOps, and Platform Engineering.

What you'll bring:

  • At least 5 years of hands-on experience as a Cloud Infrastructure Engineer, DevOps, or SRE with a proven track record of operating production cloud environments at scale.

  • You operate effectively in ambiguous, fast-changing environments. You can pick up a half-defined problem, define the path forward, and drive it to production without waiting for a playbook.

  • Cloud Platforms: Proficiency in multi-cloud operations. AWS is highly preferred; GCP is a plus.

  • Experience implementing infrastructure and security policy as code.

  • Strong software development skills, preferably in Python or another high level language

  • Strong written and verbal communication skills for driving cross-team alignment. You must be able to clearly and persuasively communicate complex concepts and risks in an engineering-driven environment.

  • Experience mentoring engineers, leading post-incident reviews, or driving cross-team infra initiatives to completion. You're comfortable being the person other engineers ask when something breaks.

  • Ability to write clean, maintainable code for automation and tooling. Experience building internal tools or services to eliminate manual work is a plus.

  • Familiarity with foundational networking concepts and protocols (TCP/IP, DNS, load balancing, VPCs, firewalls) and their application in cloud and hybrid environments.

  • Strong hands-on skills with Linux and command-line tools; you are comfortable using terminals and utilities to manage and debug systems efficiently.

  • Comfort using AI tooling as leverage for infra automation, tooling, and debugging. Bonus if you've built or contributed to AI-assisted DevOps workflows.

Our stack:

  • Cloud: AWS (EC2, IAM, VPC, ECS, Fargate, Lambda, RDS, Elasticache, Opensearch), GCP (BigQuery)

  • Monitoring/Observability: Datadog, Sentry, Amplitude

  • Github, GHA

  • IAC: Terraform, HCP

  • Security tooling

  • Codebase: Python/Typescript/React

What we offer:

  • Hybrid Work Model: Based in San Francisco with a balance of in-office and remote flexibility.

  • Fresh Lunch: Provided on in-office days.

  • Commuter Support: $150 monthly reimbursement for transit expenses.

  • Health & Wellness: $200 quarterly reimbursement to support your well-being.

  • Time Off: Flexible PTO plus 14 company holidays.

  • Comprehensive Coverage: 100% medical, dental, and vision for employees; 75% coverage for dependents.

  • Parental Leave: 16 weeks fully paid.

  • Retirement & Ownership: 401k plan plus an equity package.

  • Team Connection: Quarterly virtual events and an annual in-person summit.

Similar Jobs

9 Hours Ago
Remote or Hybrid
Mountain View, CA, USA
129K-260K Annually
Senior level
129K-260K Annually
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Designs, builds, and operates cloud-based CI/CD platforms and pipelines for embedded software production and virtualized testing. Responsibilities include automating builds, integration, testing, provisioning, monitoring, and deployment; integrating physical and virtual test environments; improving platform reliability and scalability; troubleshooting infrastructure; and guiding teams on cloud, CI/CD, virtualization, and cybersecurity best practices.
Top Skills: AutovalAWSAzureBashCC++CanCi/CdDockerDspace SystemdeskDspace VeosEmbedded Software Development ToolsEthernetGithub ActionsHardware-In-The-Loop TestingIntrepid Vehicle SpyJavaJenkinsKubernetesLinPythonSpiTest Automation FrameworksVector CanapeVector CanoeVirtualization
3 Days Ago
In-Office
San Jose, CA, USA
124K-213K Annually
Senior level
124K-213K Annually
Senior level
Fintech • Payments
Designs, builds, and operates AWS cloud networking and routing infrastructure, including VPCs, transit gateways, load balancers, CDN, WAF, ingress, and service mesh systems. Develops Terraform infrastructure as code, troubleshoots network and TLS incidents, supports edge services through on-call rotation, and documents operational practices. Partners with SRE, security, and product engineering teams while mentoring engineers and driving technical best practices.
Top Skills: AlbAWSAws CloudfrontAws Transit GatewayAws VpcBashCdnContourDatadogDnsDockerEksElbEnvoyGithub ActionsGithub EnterpriseGoKubernetesNginxNlbPythonRoute 53Tcp/IpTerraformTls/SslWaf
3 Days Ago
In-Office
Financial District, San Francisco, CA, USA
Senior level
Senior level
Information Technology
Supports and maintains AWS infrastructure in a managed services environment. Responsibilities include administering compute, storage, networking, and identity services; monitoring environments; troubleshooting incidents; performing backup, recovery, patching, and infrastructure changes; conducting root cause analysis; optimizing performance and capacity; maintaining documentation; and participating in rotational on-call support.
Top Skills: Amazon CloudwatchAmazon EbsAmazon Ec2Amazon EfsAmazon LinuxAmazon S3Amazon VpcAWSAws BackupAws CloudtrailAws IamCloudFormationDhcpDnsFirewallsInternet GatewayItilLoad BalancersNat GatewayPowershellRhelRoute TablesSecurity GroupsShellTcp/IpTerraformUbuntuVpnWindows Server

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account