Top Senior Site Reliability Engineer Jobs in San Francisco, CA

2 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
195K-258K Annually
Senior level
195K-258K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills: Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Reposted 5 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills: AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
Reposted 2 Days AgoSaved
Remote
San Francisco Bay Area, CA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 3 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 3 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
Mid level
Mid level
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills: AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
Reposted 13 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
214K-260K Annually
Senior level
214K-260K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills: AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 6 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Reposted 9 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
200K-230K Annually
Senior level
200K-230K Annually
Senior level
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills: Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Reposted 10 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
2 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills: Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Reposted 3 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 19 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted YesterdaySaved
In-Office
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Software
The Site Reliability Engineer ensures the reliability and performance of products Devin and Windsurf, managing incident response, CI/CD pipelines, infrastructure as code, and fostering a reliability culture within the engineering team.
Top Skills: AWSAzureCi/CdGCPKubernetesTerraform
6 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
147K-278K Annually
Senior level
147K-278K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
Reposted YesterdaySaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
The Site Reliability Engineer will develop, deploy, and operate AI infrastructure, focusing on high-performance and scalable machine learning systems using Kubernetes and cloud platforms.
Top Skills: AWSAzureC++GCPGoKubernetesOci
2 Days AgoSaved
In-Office
San Francisco Bay Area, CA
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Lead design, development, and operation of cloud infrastructure and SRE-focused systems. Own medium-to-large infrastructure projects, build resilient platforms, and drive cross-team technical delivery. Mentor engineers, define SLOs, reduce failure domains, and build tooling for automated, secure CI/CD and production reliability.
Top Skills: ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
Reposted 7 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
147K-278K Annually
Senior level
147K-278K Annually
Senior level
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills: AWSGoKubernetesPuppetPythonTerraform
Reposted 2 Days AgoSaved
In-Office
San Francisco Bay Area, CA
200K-275K Annually
Senior level
200K-275K Annually
Senior level
Artificial Intelligence • Healthtech • Information Technology • Software
As a Site Reliability Engineer, you will manage the production environment, focusing on infrastructure design, automation, and optimizing deployment pipelines to ensure high availability.
Top Skills: HelmKafkaKubernetesPostgresPythonRedisTerraformTypescript
Reposted 3 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
Entry level
Entry level
Artificial Intelligence • Machine Learning • Biotech • Generative AI
The Site Reliability Engineer will manage digital infrastructure, ensuring access to compute resources, automating processes, and maintaining resource visibility for researchers.
Top Skills: AnsibleDockerGrafanaKubernetesPrometheusPythonTailscaleTalos Linux
Reposted 3 Days AgoSaved
In-Office
San Francisco Bay Area, CA
200K-260K Annually
Expert/Leader
200K-260K Annually
Expert/Leader
Big Data
Lead and operate cloud infrastructure and SRE practices, focusing on IaC, CI/CD, observability, incident response, and production-grade agentic AI/LLM systems. Architect, build, and run scalable platforms, drive reliability and SLOs, author runbooks/ADRs, mentor engineers, participate in on-call rotations, and reduce toil via automation and platform tooling.
Top Skills: Agent OrchestrationAgentic AiAWSAzureDockerFluxcdGCPGithub ActionsGitopsGoGrafanaJavaJenkinsKubernetesLinuxLlmLlm GatewayLokiOpentofuPrometheusPythonSentrySignozTcp/IpTerraform
Reposted 5 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
30K-120K Annually
Senior level
30K-120K Annually
Senior level
Information Technology • Automation
The SRE/Infrastructure Engineer will architect and manage secure, scalable systems for automated penetration testing, optimizing reliability, and enhancing infrastructure based on customer demand. Responsibilities include maintaining production environments, leading technical discussions, and promoting high coding standards.
Top Skills: AWSAzureCloudFormationElkGCPNew RelicOpentelemetryPostgresPrometheusTerraform
Reposted 5 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
164K-306K Annually
Senior level
164K-306K Annually
Senior level
Software
Operate and improve reliability across Retool Cloud, managed, BYOC, and self-hosted deployments. Automate provisioning, upgrades, migrations, and secret rotations. Build observability and safer deployment/rollback workflows, partner with product teams, and produce runbooks, docs, and migration guides to reduce customer toil and scale operations.
Top Skills: AWSDocker ComposeGoHelmJavaKubernetesPostgresPythonRubyTerraformTypescript
Reposted 5 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
170K-230K Annually
Mid level
170K-230K Annually
Mid level
Artificial Intelligence • Cloud • Information Technology • Software
Contribute to the reliability and performance of Mithril's GPU orchestration platform through automation, observability, and infrastructure management. Collaborate with the team to ensure scalability across multi-cloud environments while maintaining systems stability and implementing SLOs.
Top Skills: AWSAzureGCPGoGrafanaKubernetesLinuxOpentelemetryPrometheusPulumiPythonTcp/IpTerraform
Reposted 5 Days AgoSaved
In-Office
San Francisco Bay Area, CA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
Design, operate, and automate global network and reliability infrastructure for large-scale ML workloads and a private supercomputer. Own device configuration management, protocols (BGP, VPNs, WAN), datacenter fabrics, monitoring/SLOs, incident response, security/compliance, and cross-team reliability improvements.
Top Skills: AirflowAnsibleBashBgpBluefieldCniCumulus LinuxDatadogEcmpElkEvpn/VxlanFirewallsGrafanaInfinibandInfobloxIngressIpsec VpnsIscsiKafkaKubernetesLinuxLoad BalancersLustrefsMplsNetboxNetwork PolicyNfsNornirOpentelemetryPrometheusPythonQosService NetworkingSparkSpectrum-XSpine-LeafSwitchesTerraformVpnsWan Circuits
7 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Embed with customer teams running large-scale GPU training and inference to onboard, tune, debug, and improve reliability. Diagnose fabric, driver, scheduler, and application failures; profile performance; build automation, monitoring, and preflight checks; lead incident response; and convert field learnings into product improvements and reusable reference configurations.
Top Skills: AnsibleBashCgroupsContainer RuntimesCudaDcgmDevice PluginsFabric ManagerGoGpfsHelmInfinibandKubernetesKv CacheLustreNamespacesNcclNvidia DriversNvidia-SmiNvlinkPythonRoceSlurmSshTerraformTopology-Aware SchedulingVastWeka
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account