We are searching for a Senior DevOps Engineer to take ownership of the infrastructure our platform runs on and build it forward, operating it independently, documenting it, and growing it into a true platform and developer-experience function. This is a senior, high-ownership role that is as much about building as operating: you will own domains end to end, take on the reconstruction work already scoped, and shape where the function goes next. Our platform tracks equipment across some of the most demanding construction job sites, so the infrastructure has to be reliable, observable, and cost-efficient, but keeping the lights on is the foundation you build on, not the entire job. If you have run production infrastructure on AWS and you pick up unfamiliar systems fast with sound judgment, you are exactly who we are looking for. From there, our stack is teachable, and we will get you up to speed on it. Your work here is visible and recognized, and you will have a real path to grow both the function and your own role as the company scales. Sounds like you? Apply now!
Responsibilities
- Build and evolve the platform function rather than just keep it running by taking on the reconstruction and platform work already scoped, bringing more of the environment under code and under clear ownership, and shaping how DevOps grows toward platform and developer experience.
- Own the uptime of our production platform. Lead incident response when systems degrade or fail, drive root cause analysis, and follow through with the fixes and safeguards that keep the same failure from recurring, including on-call practices, runbooks, and blameless postmortems.
- Own metrics, monitoring, alerting, and tracing across the stack so the system surfaces problems before customers feel them.
- Build dashboards and alerts that are actionable rather than noisy, and extend coverage as new services ship.
- Manage, improve, and extend our AWS environment, anchored on EKS, RDS/Aurora, and networking.
- Operate our AWS environment with sound judgment regarding scalability, redundancy, cost, and fault tolerance, and keeps it secure.
- Operate and upgrade our EKS clusters. Deploy and scale workloads, manage version upgrades, tune for reliability and cost, and troubleshoot cluster and workload issues.
- Keep our databases and messaging layer healthy and performant, including PostgreSQL on RDS/Aurora and our message broker.
- Diagnose degraded performance across data, messaging, and network paths.
- Define and manage infrastructure as code so every change is repeatable and reviewable. We build with AWS CDK in JavaScript/TypeScript; you will extend the existing stacks and bring more of the environment under code over time.
- Own the pipelines engineers rely on, primarily in GitHub Actions. Find and remove bottlenecks that slow engineering down, whether slow deploys, flaky pipelines, or scaling chokepoints, and build tooling engineers actually trust.
- Own how tracker data moves from equipment in the field through cellular tunnels and our AWS networking into the platform. Ensure pathways are reliable, secure, and observable.
- Monitor and analyze AWS spend, identify savings opportunities, and bring clear recommendations and trade-offs to engineering leadership on a regular cadence.
- Document the systems you own and distribute that knowledge across the team so the function never depends on one person.
Qualifications
- 5+ years in platform, infrastructure, DevOps, or SRE roles, including recent experience as a senior individual contributor who has owned production systems end to end.
- Substantial hands-on experience running production infrastructure on AWS.
- Hands-on production experience operating container-orchestrated workloads: deploying, scaling, upgrading, and troubleshooting them.
- Experience operating managed relational databases in production, including diagnosing and resolving performance problems.
- A solid grasp of VPCs, subnets, routing, and DNS, and how traffic actually flows between systems.
- Proficiency writing infrastructure as code and the automation around it. You write your infrastructure rather than click it.
- Experience with AWS CDK in JavaScript/TypeScript, GitHub Actions, distributed messaging, caching, or search is preferred.
- A proven ability to ramp quickly on unfamiliar systems and reach a working understanding without step-by-step direction.
- Genuine curiosity about how systems work, down to the underlying mechanics rather than treating them as black boxes.
- A strong sense of ownership: you understand, improve, and stand behind the systems in your care.
- Communicates risks, decisions, and progress proactively, and can make technical and cost trade-offs clear to people who are not infrastructure experts.
- Documents while building and shares knowledge, so no single person, including you, becomes a point of failure.
- Sound judgment about change safety: moves quickly on work that is safe and reversible, and vets larger or riskier changes before making them.
- A bachelor’s degree in a relevant field, or equivalent hands-on experience is required.
What you need to know:
- Full-time opportunity.
- Location: Fully remote - nationwide.
- Travel is required, 8-10%.
- Opportunities for growth and personal development within a highly dynamic team.
- Robust, low-cost benefit packages offered.
- Benefit coverage begins on the first date of employment.
- Paid Time Off and Volunteer Time Off offered.
- 401k match.
- Dependent Care offered.
- Employee referral bonuses.
Similar Jobs
What you need to know about the San Francisco Tech Scene
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



