Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in San Francisco Bay Area, CA
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills:
Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Cloud • Software
Operate and improve large-scale on-premises infrastructure across bare-metal servers, VMs, storage, networking, Kubernetes, containers, CI/CD, and Kafka. Automate provisioning and configuration with Ansible and related infrastructure-as-code tools, maintain observability and reliability, troubleshoot hardware and operating systems, manage disaster recovery, and support incident response. The role requires onsite work weekly at a designated data center and participation in on-call rotations.
Top Skills:
AnsibleArgo CdBare-Metal ServersBashCertificate ManagementCi/CdContainersDhcpDnsElk StackFirewallsForemanGitopsGoGrafanaHelmKafkaKubernetesLinuxLinux NetworkingLoad BalancersMaasNtpPrometheusPythonRoutingStorage AppliancesTerraformVirtual MachinesVirtualizationVlans
Artificial Intelligence • Logistics • Software
The Site Reliability Engineer will enhance operational resilience, ensuring system stability, observability, and debugging workflows for complex failures while improving developer focus and uptime.
Top Skills:
DatadogGoPrometheusPythonSentry
Artificial Intelligence • Software
As a Software Engineer on the Site Reliability team, you'll ensure system reliability, scalability, and observability while partnering with engineering teams and improving incident management processes.
Top Skills:
AWSCi/Cd ToolingContainer OrchestrationDatadogGrafanaPrometheusTerraform
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Staff Software Engineer in Site Reliability, you'll manage infrastructure for reliability and scalability, lead incident management, and automate operational tasks.
Top Skills:
AWSAzureBashCloudFormationDatadogGCPGoIncidentioPagerdutyPulumiPythonSentryTerraform
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Software Engineer in Site Reliability, you will ensure the reliability and performance of our AI platform through automation and strategic infrastructure management.
Top Skills:
AWSAzureBashCloudFormationDatadogGCPGoKubernetesPagerdutyPythonSentryTerraform
Information Technology • Software • Big Data Analytics
The Site Reliability Engineer will design, analyze, and troubleshoot large-scale distributed systems, focusing on operating systems and performance tuning.
Top Skills:
ApacheJava
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, improve workflows, manage secure infrastructure, and participate in on-call rotation for an AI-driven company.
Top Skills:
AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIpIscsiJenkinsKubernetesLinux/DebianMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRaidRubyS3ScyllaSshSslSupermicroTcpTlsUbuntu
Artificial Intelligence • Machine Learning • Robotics • Software • Transportation • Design • Manufacturing
The Staff Site Reliability Engineer will lead source control strategy, manage Git-based monorepo operations, improve developer productivity, and oversee migrations to GitHub Cloud.
Top Skills:
BazelBuckBuildkiteGerritGithub ActionsGithub CloudGithub EnterpriseGitlab CiJenkinsPulumiReviewableTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills:
ArgocdAWSDatadogGitGoKubernetesPythonTerraform
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own the reliability, scalability, and operational health of an AWS and Kubernetes-based platform supporting connected sensors. Responsibilities include production troubleshooting, incident leadership, infrastructure management with Terraform, automation, CI/CD improvements, observability, service-level objectives, runbooks, and post-incident reviews. The role partners with application, AI, embedded systems, and fleet teams and participates in on-call support.
Top Skills:
Amazon EksAWSBashCCi/CdDnsGitopsGoIamKubernetesLinuxNetworkingPythonRustTerraformVpn
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills:
AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Artificial Intelligence • Software • Generative AI
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating cloud-native systems, owning production reliability and incident response, building observability and automation, defining SLOs/SLIs, driving postmortems, and partnering with product and engineering teams to improve operational maturity.
Top Skills:
AWSAzureGCPGoJavaKubernetesPython
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills:
AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Artificial Intelligence • Big Data • Software
Own and improve infrastructure for the Data Replication platform: Kubernetes, CI/CD, secrets, networking, cloud (AWS/GCP). Drive reliability, observability, AI-augmented tooling, canary rollouts, incident reduction, runbooks, and partner with product engineers.
Top Skills:
Agentic FrameworksAirbyteAWSCdksCi/CdConnector-Based ArchitecturesDatadogGCPGrafanaHelmJavaKubernetesLlmsPrometheusPythonSecrets ManagementTerraform
Gaming
Own and operate large-scale infrastructure for sports betting and media platforms across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and alerting, support development teams, and participate in incident response. The role requires strong distributed-systems expertise, production troubleshooting, cross-team project leadership, technical communication, and mentoring.
Top Skills:
ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform
Artificial Intelligence • Healthtech
Own production reliability and performance for a healthcare AI platform: define SLIs/SLOs, run incident response and postmortems, build observability and automation, optimize scalability across serverless and container services, and partner with engineering, security, and compliance to improve operational maturity.
Top Skills:
Amazon EcsAurora PostgresAws LambdaBashClickhouseCloudwatchDatadogGithub ActionsGrafanaOpentelemetryPostgresPythonSentrySIEMVanta
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills:
Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills:
ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills:
Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills:
AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Operate and scale ThousandEyes Federal region infrastructure in a FedRAMP-compliant AWS environment. Design, deploy, and automate cloud-native services, implement IaC, monitor and audit systems, collaborate with security teams to remediate vulnerabilities, participate in 24x7 incident response and capacity planning, and ensure platform reliability, performance, and compliance.
Top Skills:
AWSFedrampGoKubernetesLinuxPuppetPythonTerraformUnixUs Govcloud
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Bay Area, CA Companies Hiring SRE Engineers
See AllPopular San Francisco Bay Area, CA Engineering Job Searches
Engineering Jobs in San Francisco Bay Area, CA
Software Engineer Jobs in San Francisco Bay Area, CA
Android Developer Jobs in San Francisco Bay Area, CA
C# Jobs in San Francisco Bay Area, CA
C++ Jobs in San Francisco Bay Area, CA
DevOps Jobs in San Francisco Bay Area, CA
Front End Developer Jobs in San Francisco Bay Area, CA
Golang Jobs in San Francisco Bay Area, CA
Hardware Engineer Jobs in San Francisco Bay Area, CA
iOS Developer Jobs in San Francisco Bay Area, CA
Java Developer Jobs in San Francisco Bay Area, CA
Javascript Jobs in San Francisco Bay Area, CA
Linux Jobs in San Francisco Bay Area, CA
Engineering Manager Jobs in San Francisco Bay Area, CA
.NET Developer Jobs in San Francisco Bay Area, CA
PHP Developer Jobs in San Francisco Bay Area, CA
Python Jobs in San Francisco Bay Area, CA
QA Jobs in San Francisco Bay Area, CA
Ruby Jobs in San Francisco Bay Area, CA
Salesforce Developer Jobs in San Francisco Bay Area, CA
Scala Jobs in San Francisco Bay Area, CA
Application Engineer Jobs in San Francisco, CA
Associate Software Engineer Jobs in San Francisco Bay Area, CA
Automation Engineer Jobs in San Francisco Bay Area, CA
AWS Engineer Jobs in San Francisco, CA
Backend Engineer Jobs in San Francisco, CA
Cloud Engineer Jobs in San Francisco Bay Area, CA
Controls Engineer Jobs in San Francisco Bay Area, CA
CTO Jobs in San Francisco Bay Area, CA
Design Engineer Jobs in San Francisco Bay Area, CA
DevOps Engineer Jobs in San Francisco Bay Area, CA
Director of Engineering Jobs in San Francisco, CA
Electrical Engineering Jobs in San Francisco Bay Area, CA
Embedded Software Engineer Jobs in San Francisco Bay Area, CA
Field Engineer Jobs in San Francisco Bay Area, CA
Firmware Engineer Jobs in San Francisco, CA
Full-Stack Engineer Jobs in San Francisco, CA
Game Engineer Jobs in San Francisco, CA
Infrastructure Engineer Jobs in San Francisco, CA
Manufacturing Engineer Jobs in San Francisco Bay Area, CA
Mechanical Design Engineer Jobs in San Francisco Bay Area, CA
Mechanical Engineering Jobs in San Francisco Bay Area, CA
Mechatronics Engineering Jobs in San Francisco Bay Area, CA
Network Engineer Jobs in San Francisco Bay Area, CA
Platform Engineer Jobs in San Francisco, CA
Process Engineer Jobs in San Francisco Bay Area, CA
Project Engineer Jobs in San Francisco Bay Area, CA
QA Engineer Jobs in San Francisco Bay Area, CA
Robotics Engineer Jobs in San Francisco Bay Area, CA
Security Engineer Jobs in San Francisco Bay Area, CA
Software Engineering Manager Jobs in San Francisco, CA
Software Test Engineer Jobs in San Francisco, CA
SRE Engineer Jobs in San Francisco Bay Area, CA
Systems Engineer Jobs in San Francisco Bay Area, CA
VP of Engineering Jobs in San Francisco, CA
All Filters
Total selected ()
No Results
No Results




















.png)












