Top SRE Engineer Jobs in San Francisco Bay Area, CA

Reposted 4 Days AgoSaved
Remote
San Francisco Bay Area, CA
136K-237K Annually
Expert/Leader
136K-237K Annually
Expert/Leader
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills: Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
5 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
76K-136K Annually
Entry level
76K-136K Annually
Entry level
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills: AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
5 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills: Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
5 Days AgoSaved
Remote
San Francisco Bay Area, CA
120K-170K Annually
Senior level
120K-170K Annually
Senior level
Cloud • Software
Operate and improve large-scale on-premises infrastructure across bare-metal servers, VMs, storage, networking, Kubernetes, containers, CI/CD, and Kafka. Automate provisioning and configuration with Ansible and related infrastructure-as-code tools, maintain observability and reliability, troubleshoot hardware and operating systems, manage disaster recovery, and support incident response. The role requires onsite work weekly at a designated data center and participation in on-call rotations.
Top Skills: AnsibleArgo CdBare-Metal ServersBashCertificate ManagementCi/CdContainersDhcpDnsElk StackFirewallsForemanGitopsGoGrafanaHelmKafkaKubernetesLinuxLinux NetworkingLoad BalancersMaasNtpPrometheusPythonRoutingStorage AppliancesTerraformVirtual MachinesVirtualizationVlans
Reposted 14 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
200K-240K Annually
Mid level
200K-240K Annually
Mid level
Artificial Intelligence • Logistics • Software
The Site Reliability Engineer will enhance operational resilience, ensuring system stability, observability, and debugging workflows for complex failures while improving developer focus and uptime.
Top Skills: DatadogGoPrometheusPythonSentry
Reposted 14 Days AgoSaved
In-Office
San Francisco Bay Area, CA
230K-390K Annually
Senior level
230K-390K Annually
Senior level
Artificial Intelligence • Software
As a Software Engineer on the Site Reliability team, you'll ensure system reliability, scalability, and observability while partnering with engineering teams and improving incident management processes.
Top Skills: AWSCi/Cd ToolingContainer OrchestrationDatadogGrafanaPrometheusTerraform
Reposted 14 Days AgoSaved
In-Office
San Francisco Bay Area, CA
238K-290K Annually
Expert/Leader
238K-290K Annually
Expert/Leader
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Staff Software Engineer in Site Reliability, you'll manage infrastructure for reliability and scalability, lead incident management, and automate operational tasks.
Top Skills: AWSAzureBashCloudFormationDatadogGCPGoIncidentioPagerdutyPulumiPythonSentryTerraform
Reposted 14 Days AgoSaved
In-Office
San Francisco Bay Area, CA
200K-260K Annually
Mid level
200K-260K Annually
Mid level
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Software Engineer in Site Reliability, you will ensure the reliability and performance of our AI platform through automation and strategic infrastructure management.
Top Skills: AWSAzureBashCloudFormationDatadogGCPGoKubernetesPagerdutyPythonSentryTerraform
Reposted 14 Days AgoSaved
In-Office
San Francisco Bay Area, CA
Mid level
Mid level
Information Technology • Software • Big Data Analytics
The Site Reliability Engineer will design, analyze, and troubleshoot large-scale distributed systems, focusing on operating systems and performance tuning.
Top Skills: ApacheJava
Reposted 14 Days AgoSaved
In-Office
San Francisco Bay Area, CA
160K-250K Annually
Mid level
160K-250K Annually
Mid level
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, improve workflows, manage secure infrastructure, and participate in on-call rotation for an AI-driven company.
Top Skills: AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIpIscsiJenkinsKubernetesLinux/DebianMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRaidRubyS3ScyllaSshSslSupermicroTcpTlsUbuntu
Reposted 14 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
230K-315K Annually
Senior level
230K-315K Annually
Senior level
Artificial Intelligence • Machine Learning • Robotics • Software • Transportation • Design • Manufacturing
The Staff Site Reliability Engineer will lead source control strategy, manage Git-based monorepo operations, improve developer productivity, and oversee migrations to GitHub Cloud.
Top Skills: BazelBuckBuildkiteGerritGithub ActionsGithub CloudGithub EnterpriseGitlab CiJenkinsPulumiReviewableTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 14 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills: ArgocdAWSDatadogGitGoKubernetesPythonTerraform
15 Days AgoSaved
In-Office
San Francisco Bay Area, CA
Entry level
Entry level
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own the reliability, scalability, and operational health of an AWS and Kubernetes-based platform supporting connected sensors. Responsibilities include production troubleshooting, incident leadership, infrastructure management with Terraform, automation, CI/CD improvements, observability, service-level objectives, runbooks, and post-incident reviews. The role partners with application, AI, embedded systems, and fleet teams and participates in on-call support.
Top Skills: Amazon EksAWSBashCCi/CdDnsGitopsGoIamKubernetesLinuxNetworkingPythonRustTerraformVpn
Reposted 16 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Reposted 16 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Software • Generative AI
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating cloud-native systems, owning production reliability and incident response, building observability and automation, defining SLOs/SLIs, driving postmortems, and partnering with product and engineering teams to improve operational maturity.
Top Skills: AWSAzureGCPGoJavaKubernetesPython
Reposted 16 Days AgoSaved
In-Office
San Francisco Bay Area, CA
131K-192K Annually
Senior level
131K-192K Annually
Senior level
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills: AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Reposted 16 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
196K-235K Annually
Senior level
196K-235K Annually
Senior level
Artificial Intelligence • Big Data • Software
Own and improve infrastructure for the Data Replication platform: Kubernetes, CI/CD, secrets, networking, cloud (AWS/GCP). Drive reliability, observability, AI-augmented tooling, canary rollouts, incident reduction, runbooks, and partner with product engineers.
Top Skills: Agentic FrameworksAirbyteAWSCdksCi/CdConnector-Based ArchitecturesDatadogGCPGrafanaHelmJavaKubernetesLlmsPrometheusPythonSecrets ManagementTerraform
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
145K-193K Annually
Senior level
145K-193K Annually
Senior level
Gaming
Own and operate large-scale infrastructure for sports betting and media platforms across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and alerting, support development teams, and participate in incident response. The role requires strong distributed-systems expertise, production troubleshooting, cross-team project leadership, technical communication, and mentoring.
Top Skills: ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform
Reposted 17 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Healthtech
Own production reliability and performance for a healthcare AI platform: define SLIs/SLOs, run incident response and postmortems, build observability and automation, optimize scalability across serverless and container services, and partner with engineering, security, and compliance to improve operational maturity.
Top Skills: Amazon EcsAurora PostgresAws LambdaBashClickhouseCloudwatchDatadogGithub ActionsGrafanaOpentelemetryPostgresPythonSentrySIEMVanta
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
90K-100K Annually
Mid level
90K-100K Annually
Mid level
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills: Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
125K-150K Annually
Mid level
125K-150K Annually
Mid level
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills: ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
9 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
136K-181K Annually
Entry level
136K-181K Annually
Entry level
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills: Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
Reposted 18 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills: AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Reposted 19 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Operate and scale ThousandEyes Federal region infrastructure in a FedRAMP-compliant AWS environment. Design, deploy, and automate cloud-native services, implement IaC, monitor and audit systems, collaborate with security teams to remediate vulnerabilities, participate in 24x7 incident response and capacity planning, and ensure platform reliability, performance, and compliance.
Top Skills: AWSFedrampGoKubernetesLinuxPuppetPythonTerraformUnixUs Govcloud
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account