Top SRE Engineer Jobs in San Francisco Bay Area, CA

8 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
165K-227K Annually
Senior level
165K-227K Annually
Senior level
Cloud
Design, operate, and improve highly available cloud services and infrastructure in a FedRAMP environment. Build automation and internal platforms using Go, Python, Terraform, Kubernetes, and GitOps practices. Participate in global on-call and incident response, define SLIs and SLOs, improve observability, and drive post-incident reliability improvements. Collaborate across engineering teams, provide technical guidance and mentoring, support compliance readiness, and own reliability initiatives from design through production.
Top Skills: Amazon EksArgocdAWSCassandraCi/CdDatadogFedrampGCPGitGitopsGoGoogle GkeGrafanaHelmHipaaKubernetesLinuxMySQLOpensearchPostgresPythonRedisRustSnowflakeSoc 2SplunkTerraformTerraform
8 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills: AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
Reposted 9 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
200K-240K Annually
Senior level
200K-240K Annually
Senior level
Artificial Intelligence • Software
Design, build, and scale control- and data-plane infrastructure for distributed AI workloads. Improve reliability, performance, scheduling, and observability for Ray clusters across cloud and on-prem environments. Support accelerator integration, container image management, and provide on-call troubleshooting and cross-team collaboration.
Top Skills: AWSAzureContainersGCPGoGpusGrafanaKubernetesLinuxPrometheusPythonRayTpusVms
10 Days AgoSaved
In-Office
San Francisco Bay Area, CA
121K-151K Annually
Senior level
121K-151K Annually
Senior level
Fintech • Payments
Leads large-scale site reliability engineering strategy, architecting highly available and scalable systems, improving observability, automation, incident response, capacity planning, performance, and cloud costs. Builds self-healing mechanisms and AI agents that automate operational workflows, reduce TOIL, and support incident response and anomaly detection. Establishes AI security and governance controls, advises engineering leadership, leads cross-functional reliability initiatives, and mentors engineers developing production-grade SRE and agentic solutions.
Top Skills: APIsCi/CdDistributed TracingDockerElk StackGrafanaJaegerKubernetesMySQLNoSQLOpentelemetryPostgresPrometheusService MeshesSplunk
Reposted 2 Hours AgoSaved
Remote
San Francisco Bay Area, CA
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Reposted 11 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills: AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
2 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills: Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
2 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills: DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Reposted 11 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
140K-288K Annually
Senior level
140K-288K Annually
Senior level
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills: ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Reposted 11 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
Reposted 11 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills: AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Reposted 3 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
138K-221K Annually
Senior level
138K-221K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills: .NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
3 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills: Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Reposted 3 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills: ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Reposted 12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
116K-200K Annually
Mid level
116K-200K Annually
Mid level
Information Technology • Mobile • Software
As a Site Reliability Engineer, you'll ensure system reliability and scalability, automate processes, optimize performance, and collaborate on system design.
Top Skills: AWSAzureBashCloudFormationDatadogDockerElkGoGoogle Cloud PlatformGrafanaHelmKubernetesNew RelicPrometheusPulumiPythonTerraform
Reposted 12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
230K-310K Annually
Senior level
230K-310K Annually
Senior level
Artificial Intelligence • Software
Own the reliability and performance of backend systems at Gamma, building automation and tooling while leading incident response and improving system stability.
Top Skills: AWSCloudFormationDockerGoKafkaKubernetesNode.jsPythonTerraformTypescript
Reposted 12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
Mid level
Mid level
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own reliability for a live fleet of Linux-based edge sensors and cloud infrastructure. Triage and recover field hardware, perform SSH-based diagnostics, build fleet management and OTA systems, implement observability and alerting, automate operational tasks, develop runbooks, and participate in on-call rotations to prevent and resolve incidents.
Top Skills: AWSBashCDnsDockerFirewallsGoIamKubernetesLinuxPythonRustSshVpn
Reposted 12 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
140K-210K Annually
Junior
140K-210K Annually
Junior
Artificial Intelligence • Hardware • Robotics • Software
As a Software Engineer - Infrastructure at Skydio, you'll support and enhance Kubernetes infrastructure, improve continuous delivery, and collaborate on security and architectural improvements.
Top Skills: Cloud PlatformsContinuous DeploymentGoInfrastructureKubernetesPython
Reposted 12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
The Senior Site Reliability Engineer will design and maintain scalable infrastructure, improve system reliability, manage CI/CD pipelines, and collaborate across teams for operational excellence.
Top Skills: AnsibleArgocdAWSBashDatadogDockerElkGithub ActionsGrafanaKubernetesLinuxOpentelemetryPrometheusPythonTerraform
Reposted 12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
135K-159K Annually
Senior level
135K-159K Annually
Senior level
Big Data • Cloud • Marketing Tech • Social Impact • Software
The Senior Site Reliability Engineer will support global product deployments, provide 24/7 operational support, maintain CI/CD tooling, and optimize system performance. They will utilize their SRE practices and leadership abilities to improve product reliability and guide other engineers.
Top Skills: AWSCassandraCircleCIDynamoDBGCPGoJenkinsKubernetesNosql DatabasesPythonScylladbSinglestore DbTerraform
Reposted 12 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
167K-197K Annually
Senior level
167K-197K Annually
Senior level
Big Data • Cloud • Marketing Tech • Social Impact • Software
As a Senior Site Reliability Engineer, you will manage global product deployments, provide operational support, enhance CI/CD tooling, and optimize system performance, collaborating closely with distributed engineering teams.
Top Skills: AWSCassandraCircleCIDynamoDBGCPGoJenkinsKubernetesNosql DatabasesPythonScylladbSinglestore DbTerraform
Reposted One Month AgoSaved
Remote
San Francisco Bay Area, CA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
3 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills: AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
3 Days AgoSaved
Remote
San Francisco Bay Area, CA
115K-175K Annually
Entry level
115K-175K Annually
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
4 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills: AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account