Customer.io Logo

Customer.io

Senior Site Reliability Engineer

Sorry, this job was removed at 09:09 a.m. (PST) on Wednesday, May 20, 2026
Remote
Hiring Remotely in United States
Remote
Hiring Remotely in United States

Similar Jobs

19 Hours Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Yesterday
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
2 Days Ago
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills: AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
About Customer.io

Over 8,000 companies - from scrappy startups to global brands - use our platform to send billions of emails, push notifications, in-app messages, and SMS every day. Customer.io powers automated communication that people actually want to receive. 

We help teams send smarter, more relevant messages using real-time behavioral data. Under the hood: Go, React, Ember and AI help us ship fast and scale with confidence.

We’re looking for a Site Reliability Engineer to help us scale our infrastructure, reduce operational toil, and increase reliability as we grow. If you’ve worked on high-scale systems and love making platforms better for developers and customers alike, we’d love to meet you.

What We Value

Ownership
You own problems end to end. You move fast, act like an owner, and thrive in ambiguity. You've led complex projects before, whether officially or not, and you're ready to do it again.

Engineers with product taste

You think like a user, not just an engineer. You think about performance, reliability, and how systems impact the customer experience.

A healthy skepticism for “the way things are done”
You bring rigor and creativity. Best practices matter - but never more than forward motion.

What You’ll Do
  • Build and scale infrastructure to support billions of messages per day and real-time events
  • Automate deployments, alerting, and incident response
  • Make our on-call better - clear alerts, solid documentation, and faster resolution
  • Tune MySQL and other datastore performance and improve reliability across distributed systems
  • Collaborate across teams to debug, ship, and support systems in production
  • Share knowledge and raise the bar through sharing your progress publicly with short videos, thoughtful writing, and mentorship
  • Leverage AI tools to prototype, move faster, and make better decisions
What we're looking for
  • 7+ years in SRE or infrastructure roles, improving production systems at scale
  • Deep MySQL experience - schema design, performance tuning, and operational tooling
  • Fluency in cloud-native tech (GCP a plus) and Terraform
  • Proficiency in Go and Bash for scripting and systems programming
  • Skill in observability, incident response, and debugging distributed systems
  • A preference for action over perfection, and pride in owning technical decisions
Compensation & Benefits

We believe in transparency. Starting salary for this role is $140,000 - $180,000 USD (or equivalent in local currency) depending on experience and subject to market rate adjustment.

We know our people are what make us great, and we’re committed to taking great care of them. Our inclusive benefits package supports your well-being and growth, including 100% coverage of medical, dental, vision, mental health, and supplemental insurance premiums for you and your family. We also offer 16 weeks paid parental leave, unlimited PTO, stipends for remote work and wellness, a professional development budget, and more.

See full benefits here →

Our Process

No gotchas, no trick questions - just a clear, human process designed to help both of us make an informed decision.

  • Application - We review everyone with care. Tell us why you're interested.
  • Recruiter Call (30 mins) - Let’s chat about what you’re looking for and how we work.
  • Behavioral Interview (60 mins) - Talk with one of our hiring managers. We’ll explore topics like ownership, product thinking, and collaboration.
  • Take-Home Assignment - Complete a short, realistic task similar to what you’d work on here.
  • Technical Interview + Assignment Review Call (90 mins) - Walk us through your take-home project and the decisions you made along the way. We’ll also collaborate on a system design problem, focusing on real-world scaling challenges and tradeoffs.

All final candidates will be asked to complete a background check and employment verifications as part of our pre-employment process.

Customer.io recognizes the stifling impact of systemic injustice on diverse communities. We commit to using our influence to increase inclusion and equity within the tech industry. We strive to build an inclusive team culture, implement bias-free hiring practices, and develop community partnerships to expand our global impact.

Zoom is the only video conference platform that we use, virtual interviews will be conducted using the video capability (i.e., not via the chat), and offers will be extended in writing on official Customer.io letterhead. Please be vigilant in all of your job search activity, and if you have any questions please contact [email protected].

Join us!

We believe in empathy, transparency, responsibility, and, yes, a little awkwardness. If you’re excited by what you read and want to build software that makes communication better for everyone—apply now.

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account