Gamma (gamma.app) Logo

Gamma (gamma.app)

Site Reliability Engineer

Reposted 9 Days Ago
In-Office
San Francisco, CA, USA
230K-310K Annually
Senior level
In-Office
San Francisco, CA, USA
230K-310K Annually
Senior level
Own the reliability and performance of backend systems at Gamma, building automation and tooling while leading incident response and improving system stability.
The summary above was generated by AI
About the role

Gamma's infrastructure needs to be rock-solid for millions of daily users while enabling our engineering teams to ship fast. You'll own the operational health of our full backend platform, building automation and tooling that improves reliability and partnering with engineering to design systems that are observable, resilient, and easy to operate. Your work directly impacts every Gamma user's experience.

This is a high-impact role where you'll balance reliability with velocity, knowing when to move fast and when to prioritize stability. You'll lead incident response, drive systemic improvements, and help shape how Gamma scales to serve its next 100 million users.

Our team has a strong in-office culture and works in person 4–5 days per week in San Francisco. We love working together to stay creative and connected, with flexibility to work from home when focus matters most.

What you'll do
  • Own the reliability, availability, and performance of Gamma's production systems across our AWS infrastructure

  • Build observability infrastructure from the ground up: metrics, logging, tracing, and alerting that give the team genuine visibility into system health before users feel the impact

  • Design and ship automation that reduces toil, makes deployments safer, and gets us back on our feet faster when things go wrong

  • Lead incident response and blameless post-mortems, then follow through on the systemic fixes that keep the same issues from coming back

  • Partner with engineering teams on architecture reviews, SLO and SLI design, and reliability best practices that scale with the product

  • Manage and optimize our compute, networking, databases, and managed services

What you'll bring
  • 5+ years in site reliability engineering, DevOps, or systems engineering with deep, hands-on AWS expertise

  • Strong programming skills in Python, Go, or TypeScript/Node.js, applied to building real tools and automation

  • Solid experience with infrastructure-as-code (Terraform, CloudFormation) and end-to-end observability solutions

  • Track record of making systems meaningfully more reliable through automation, smarter monitoring, and architectural improvements

  • Deep understanding of networking, distributed systems, containerization (Docker, Kubernetes), and database performance at scale

  • Sharp incident management instincts and the debugging skills to navigate complex production failures

  • Experience scaling SaaS products to millions of users, or background with Kafka, chaos engineering, or service mesh technologies (Nice to have)

  • AWS certifications, or experience with security and compliance frameworks like SOC 2 or ISO 27001 (Nice to have)

Compensation range:

The base salary for this full-time position, which spans multiple internal levels depending on qualifications, ranges between $230K - $310K plus benefits & equity.

Final offer amounts are determined by multiple factors, including but not limited to experience and expertise in the requirements listed above.

If you're interested in this role but you don't meet every requirement, we encourage you to apply anyway! We're always excited about meeting great people.

HQ

Gamma (gamma.app) San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

7 Days Ago
Hybrid
San Francisco, CA, USA
214K-260K Annually
Senior level
214K-260K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills: AWSAzureDockerGCPJavaKubernetesLinuxTerraform
8 Days Ago
Hybrid
133K-226K Annually
Senior level
133K-226K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
As a Site Reliability Engineer, you will deploy and monitor cloud solutions, implement automation and infrastructure support, and collaborate across teams to enhance service delivery and performance.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsKubernetesNfsObject StorageOpen StackPaasSaaSVMware
10 Days Ago
Easy Apply
Hybrid
San Jose, CA, USA
Easy Apply
119K-170K Annually
Senior level
119K-170K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
The role involves creating scalable solutions using Linux and Kubernetes, troubleshooting performance issues, maintaining security, and writing automation tools.
Top Skills: AnsibleBashDockerFirewall TechnologiesGoKubernetesKvmLinuxMulti-Factor AuthenticationOpenstackPgpPkiPythonSshUnix

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account