Top Senior Site Reliability Engineer Jobs in San Francisco, CA

19 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills: DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Reposted 28 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
175K-285K Annually
Mid level
175K-285K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
Reposted 28 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills: AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Reposted 29 Days AgoSaved
In-Office
San Francisco Bay Area, CA
116K-200K Annually
Mid level
116K-200K Annually
Mid level
Information Technology • Mobile • Software
As a Site Reliability Engineer, you'll ensure system reliability and scalability, automate processes, optimize performance, and collaborate on system design.
Top Skills: AWSAzureBashCloudFormationDatadogDockerElkGoGoogle Cloud PlatformGrafanaHelmKubernetesNew RelicPrometheusPulumiPythonTerraform
Reposted 29 Days AgoSaved
In-Office
San Francisco Bay Area, CA
230K-310K Annually
Senior level
230K-310K Annually
Senior level
Artificial Intelligence • Software
Own the reliability and performance of backend systems at Gamma, building automation and tooling while leading incident response and improving system stability.
Top Skills: AWSCloudFormationDockerGoKafkaKubernetesNode.jsPythonTerraformTypescript
Reposted 29 Days AgoSaved
In-Office
San Francisco Bay Area, CA
Mid level
Mid level
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own reliability for a live fleet of Linux-based edge sensors and cloud infrastructure. Triage and recover field hardware, perform SSH-based diagnostics, build fleet management and OTA systems, implement observability and alerting, automate operational tasks, develop runbooks, and participate in on-call rotations to prevent and resolve incidents.
Top Skills: AWSBashCDnsDockerFirewallsGoIamKubernetesLinuxPythonRustSshVpn
Reposted 29 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
140K-210K Annually
Junior
140K-210K Annually
Junior
Artificial Intelligence • Hardware • Robotics • Software
As a Software Engineer - Infrastructure at Skydio, you'll support and enhance Kubernetes infrastructure, improve continuous delivery, and collaborate on security and architectural improvements.
Top Skills: Cloud PlatformsContinuous DeploymentGoInfrastructureKubernetesPython
Reposted One Month AgoSaved
Remote
San Francisco Bay Area, CA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
20 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills: AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
20 Days AgoSaved
Remote
San Francisco Bay Area, CA
115K-175K Annually
Entry level
115K-175K Annually
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Reposted One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
21 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills: AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
8 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
180K-230K Annually
Senior level
180K-230K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Automation
Own the reliability, scalability, and observability of production systems. Design and operate AWS infrastructure with Terraform and Kubernetes, build monitoring and SLOs, automate deployments and capacity management, ensure database and data pipeline reliability, partner on architecture reviews, and drive infrastructure security, compliance, and incident response.
Top Skills: AWSBashCi/CdCloud PlatformsDatadogDistributed SystemsGithub ActionsGrafanaInfrastructure As CodeKubernetesPrometheusPythonTerraform
One Month AgoSaved
In-Office
San Francisco Bay Area, CA
170K-225K Annually
Senior level
170K-225K Annually
Senior level
Fintech • Software • Financial Services
Own AWS and Kubernetes reliability, deployments, observability, incident response, infrastructure as code, CI/CD, and compliance controls for SOC 2 and PCI-DSS. Build scalable platform foundations, AI infrastructure, and greenfield product infrastructure while improving cost efficiency, security, vendor flexibility, and deployment consistency. Participate in on-call support and partner with security and engineering teams across backend, mobile, data, ML, and AI.
Top Skills: Ai/Llm ToolingAmazon EksAWSAws CdkCi/CdDatadogKotlinKubernetesMcpNode.jsOpentelemetryOpentofuPci-DssPythonSoc 2TerraformTypescript
22 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
76K-136K Annually
Entry level
76K-136K Annually
Entry level
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills: AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
22 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills: Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Reposted One Month AgoSaved
Hybrid
San Francisco Bay Area, CA
200K-240K Annually
Mid level
200K-240K Annually
Mid level
Artificial Intelligence • Logistics • Software
The Site Reliability Engineer will enhance operational resilience, ensuring system stability, observability, and debugging workflows for complex failures while improving developer focus and uptime.
Top Skills: DatadogGoPrometheusPythonSentry
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
230K-390K Annually
Senior level
230K-390K Annually
Senior level
Artificial Intelligence • Software
As a Software Engineer on the Site Reliability team, you'll ensure system reliability, scalability, and observability while partnering with engineering teams and improving incident management processes.
Top Skills: AWSCi/Cd ToolingContainer OrchestrationDatadogGrafanaPrometheusTerraform
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
238K-290K Annually
Expert/Leader
238K-290K Annually
Expert/Leader
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Staff Software Engineer in Site Reliability, you'll manage infrastructure for reliability and scalability, lead incident management, and automate operational tasks.
Top Skills: AWSAzureBashCloudFormationDatadogGCPGoIncidentioPagerdutyPulumiPythonSentryTerraform
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
200K-260K Annually
Mid level
200K-260K Annually
Mid level
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Software Engineer in Site Reliability, you will ensure the reliability and performance of our AI platform through automation and strategic infrastructure management.
Top Skills: AWSAzureBashCloudFormationDatadogGCPGoKubernetesPagerdutyPythonSentryTerraform
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
Mid level
Mid level
Information Technology • Software • Big Data Analytics
The Site Reliability Engineer will design, analyze, and troubleshoot large-scale distributed systems, focusing on operating systems and performance tuning.
Top Skills: ApacheJava
Reposted One Month AgoSaved
Hybrid
San Francisco Bay Area, CA
195K-265K Annually
Senior level
195K-265K Annually
Senior level
Artificial Intelligence • Machine Learning • Robotics • Software • Transportation • Design • Manufacturing
The Staff Site Reliability Engineer will lead source control strategy, manage Git-based monorepo operations, improve developer productivity, and oversee migrations to GitHub Cloud.
Top Skills: BazelBuckBuildkiteGerritGithub ActionsGithub CloudGithub EnterpriseGitlab CiJenkinsPulumiReviewableTerraform
Reposted One Month AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills: ArgocdAWSDatadogGitGoKubernetesPythonTerraform
Reposted One Month AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
One Month AgoSaved
In-Office
San Francisco Bay Area, CA
Entry level
Entry level
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own the reliability, scalability, and operational health of an AWS and Kubernetes-based platform supporting connected sensors. Responsibilities include production troubleshooting, incident leadership, infrastructure management with Terraform, automation, CI/CD improvements, observability, service-level objectives, runbooks, and post-incident reviews. The role partners with application, AI, embedded systems, and fleet teams and participates in on-call support.
Top Skills: Amazon EksAWSBashCCi/CdDnsGitopsGoIamKubernetesLinuxNetworkingPythonRustTerraformVpn
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account