Stuut Jobs

Lead Site Reliability Engineer

Stuut

Lead Site Reliability Engineer

Reposted 22 Days Ago

Be an Early Applicant

In-Office

San Francisco, CA, USA

200K-275K Annually

Senior level

In-Office

San Francisco, CA, USA

200K-275K Annually

Senior level

Lead SRE responsible for reliability strategy, architecting resilient AWS/Kubernetes infrastructure, building observability, driving incident response and postmortems, improving deployment safety and automation, mentoring engineers, and partnering across product and engineering to scale platform reliability.

The summary above was generated by AI

Stuut is transforming accounts receivable for B2B companies—making collections smarter and faster for companies that have historically relied on manual processes that are labor intensive and costly. Our platform is gaining traction with finance teams across industrials, chemicals, and manufacturing sectors from Fortune 10 brands to scaling midmarkets. We're backed by top-tier investors including a16z, Khosla, Activant, 1984 Ventures and Page One.

The Role
We’re hiring a Lead Site Reliability Engineer to drive the strategy, architecture, and execution of reliability, scalability, and operational excellence across our platform. You’ll build and scale the systems that keep Stuut highly available, performant, and resilient as we grow customers, traffic, and complexity.

From defining SLOs and reliability standards to hardening infrastructure, improving observability, and guiding teams through incident response and postmortems, you’ll own the engineering rigor that allows us to ship quickly without sacrificing stability. You’ll turn strong reliability engineering into real customer trust — creating the guardrails that let product and engineering move fast with confidence.

This is a hands-on technical leadership role for an engineer who excels at designing reliable distributed systems, influencing engineering practices, and leading high-impact reliability initiatives across teams.

What You’ll Do

Set the Reliability Strategy: define the long-term vision for site reliability, including SLOs/SLIs, error budgets, availability targets, and operational standards.
Build & Scale Reliable Infrastructure: architect and maintain resilient, scalable cloud infrastructure across AWS and Kubernetes, ensuring systems are secure, fault-tolerant, and cost-effective.
Own Observability & Monitoring: design and evolve monitoring, alerting, and logging systems that provide clear, actionable signals across services and environments.
Lead Incident Response & Postmortems: own incident management practices, lead major incident response, and drive blameless postmortems that result in meaningful system improvements.
Improve System Resilience: identify reliability risks and lead efforts around redundancy, failover, capacity planning, and graceful degradation.
Optimize CI/CD & Deployment Reliability: partner with engineering teams to ensure deployments are safe, observable, and reversible; improve rollout strategies and reduce operational risk.
Partner with Product & Engineering Teams: collaborate early in the development lifecycle to influence system design, scalability, and reliability tradeoffs.
Reduce Toil & Improve Developer Experience: automate operational tasks, improve runbooks, and build tooling that reduces manual work and accelerates safe execution.
Drive Root Cause Resolution: guide teams through deep debugging of reliability issues, ensuring fixes address underlying causes rather than symptoms.
Influence Reliability Culture: promote reliability-first thinking, strong operational hygiene, and shared ownership of production systems across engineering.
Mentor & Level Up the Team: coach engineers on reliability principles, incident handling, infrastructure design, and operational best practices.

You Might Be a Fit If You…

Have 7+ years of experience in site reliability engineering, infrastructure engineering, or backend software engineering.
Have designed and operated highly available, production-grade systems supporting rapid product iteration.
Are fluent in Python and/or TypeScript, and comfortable building automation and tooling to support reliability goals.
Have a deep experience with AWS, Kubernetes (EKS), Docker, and cloud-native architectures.
Have implemented and evolved observability stacks (metrics, logs, traces) and know how to create high-signal alerting.
Understand how to design, measure, and enforce SLOs, SLIs, and error budgets.
Have supported systems built with modern stacks such as FastAPI, Vue.js, PostgreSQL (RDS), and event-driven architectures.
Have improved reliability and operational maturity in environments using CI/CD pipelines, infrastructure as code, and modern deployment workflows.
Can balance reliability, velocity, and cost — making pragmatic tradeoffs that serve customers and the business.
Enjoy collaborating across Product, Backend, Frontend, and Infrastructure teams to improve system health.
Thrive in a role that blends deep technical execution, system design, and leadership influence in a fast-moving environment.

Compensation

Top-of-market salary and equity package
Benefits (for U.S.-based full-time employees)
Medical, dental & vision insurance coverage for you
401(k) & Match
Equity
Flexible PTO
Parental Leave

Similar Jobs

JPMorganChase

Site Reliability Engineer

4 Days Ago

Hybrid

Palo Alto, CA, USA

Senior level

Financial Services

Lead design and implementation of SRE practices, observability, reliability and AI-assisted operational workflows. Mentor engineers, define NFRs/SLOs, build automation, logging/metrics/tracing pipelines, containerized CI/CD/GitOps, and integrate AI agents for production-grade reliability.

Top Skills: AutogenCassandraChaos MonkeyChromaClaudeCrewaiDatadogDockerDynamoDBDynatraceFlinkFluentdGithub CopilotGitopsGo (Golang)GrafanaGremlinHadoopInfluxdbJavaKafkaKubernetesLangchainLanggraphLitmuschaosLogstashModel Context Protocol (Mcp)MongoDBNeo4JPineconePrometheusPythonPyTorchRabbitMQScikit-LearnSparkSplunkSqsTensorFlowTerraformTigergraphTimescaledbVectorWeaviate

JPMorganChase

Site Reliability Engineer

4 Days Ago

Hybrid

Palo Alto, CA, USA

Senior level

Financial Services

Lead design and delivery of reliability, observability, and SRE practices for large-scale data and AI/ML platforms. Define NFRs, SLI/SLOs, incident response, and implement reliable, secure, scalable infrastructure and automation while mentoring teams and driving safe, auditable AI-assisted operations.

Top Skills: AWSAws GlueCi/CdDatabricksDatadogDockerDynatraceGrafanaKubernetesMapreducePrometheusPythonSparkSplunkTerraform

Green Dot Corporation

Site Reliability Engineer

11 Days Ago

In-Office

140K-199K Annually

Senior level

140K-199K Annually

Senior level

Fintech • Financial Services

Lead Site Reliability Engineer responsible for ensuring system reliability, scalability, and performance. Develop automated deployment strategies, maintain monitoring/observability, define SLIs/SLOs, collaborate with cross-functional teams, drive reliability best practices, participate in on-call incident response, and improve delivery through automation and training.

Top Skills: AWSAzureBashGCPPowershellPython

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine