Top Reliability Engineer Jobs in San Francisco, CA

7 Days AgoSaved
In-Office
San Francisco Bay Area, CA
150K-170K Annually
Mid level
150K-170K Annually
Mid level
Artificial Intelligence • Logistics • Robotics • Software
Own reliability across cloud, edge, and on-site deployments. Build observability, monitoring, and alerting. Define incident response and on-call processes, improve deployment workflows, diagnose infra/network/distributed-system issues, and make deployments repeatable and scalable.
Top Skills: AWSAzureGCPGrafanaKafkaKubernetesLinuxOpentelemetryPrometheusRtspSecure TunnelsVpnWebrtc
Reposted 3 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 3 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
Mid level
Mid level
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills: AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
Reposted 8 Days AgoSaved
In-Office
San Francisco Bay Area, CA
160K-220K Annually
Senior level
160K-220K Annually
Senior level
Cloud
The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.
Top Skills: AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform
Reposted 13 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
214K-260K Annually
Senior level
214K-260K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills: AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 13 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 9 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Information Technology • Software
Lead end-to-end platform reliability: define SLIs/SLOs, harden production architecture, ensure Kubernetes runtime and queue safety, run incident command for Sev1/Sev2, own observability/on-call/runbooks, and gate risky releases while delivering a prioritized reliability roadmap.
Top Skills: BullmqKoaKubernetesNode.jsPostgraphilePostgresReactRedisTypescript
Reposted 9 Days AgoSaved
In-Office
San Francisco Bay Area, CA
175K-215K Annually
Junior
175K-215K Annually
Junior
Automotive
Software Reliability Engineers at Waymo ensure the stable operation of autonomous systems, collaborating on reliability solutions and system performance improvements.
Top Skills: C++JavaPython
Reposted 9 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
150K-180K Annually
Senior level
150K-180K Annually
Senior level
Aerospace • Automation
The Senior Reliability Engineer will establish AeroVect's reliability engineering practice, leading reliability analyses and managing external testing programs to ensure product durability and performance across operational environments.
Top Skills: Accelerated Life TestingData AnalysisEnvironmental TestingFmeaFtaMtbfMttrRbd
Reposted 6 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Reposted YesterdaySaved
Remote or Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Software
Lead reliability engineering for Silicon Photonics hardware: define and validate reliability models, perform MTBF/MTBCF predictions, analyze field data, direct verification testing and root-cause analysis, drive corrective actions, and mentor cross-functional teams to improve product reliability.
Top Skills: Derating AnalysisDfmeaMtbcfMtbfSherlockSilicon PhotonicsTelcordiaThermal DesignWindchill Qs
Reposted YesterdaySaved
Remote
San Francisco Bay Area, CA
75K-150K Annually
Senior level
75K-150K Annually
Senior level
Database • Analytics
As a Database Reliability Engineer at ClickHouse, you'll improve reliability, manage escalation processes, support incident response, and enhance database performance while collaborating across teams.
Top Skills: AWSAzureC++ClickhouseGoogle Cloud PlatformPythonShellSQL
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
2 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Junior
Junior
Software
Drive reliability qualification and production monitoring for optical communication products. Track and analyze reliability stress tests, investigate failures with cross-functional teams, implement corrective actions, maintain dashboards and reports, support NPI, and apply statistical and AI tools to generate reliability insights and improve product quality.
Top Skills: AIJmp/JslMinitabSQL
Reposted 11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
105K-137K Annually
Senior level
105K-137K Annually
Senior level
Food • Marketing Tech • Manufacturing
The Senior Reliability Engineer enhances equipment reliability, reduces downtime, and improves maintenance strategies across production. This role involves collaboration with engineering and operations, leading reliability programs, and mentoring junior staff.
Top Skills: Advanced AnalyticsCmmsDigital ToolsReliability Modeling Tools
Reposted 11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
180K-230K Annually
Senior level
180K-230K Annually
Senior level
Energy • Renewable Energy
The Staff Reliability Engineer will ensure hardware reliability in high-voltage electronics, develop reliability test programs, and collaborate on design and testing across teams.
Top Skills: Hv ElectronicsPower ConversionPython
Reposted 7 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
227K-272K Annually
Senior level
227K-272K Annually
Senior level
eCommerce • Healthtech • Kids + Family • Retail • Social Media
Own and evolve Babylist's AWS infrastructure and developer platform using Terraform and Kubernetes. Improve CI/CD reliability, support engineers across environments, define monitoring and alerting standards, lead incident response and postmortems, and shape platform architecture to scale for millions of users.
Top Skills: AWSCdnCircleCICronitorDatadogDnsEksGithub ActionsKubernetesLoad BalancersMySQLPagerdutyRdsRedisRuby On RailsSentrySidekiqTerraform
Reposted 11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
150K-180K Annually
Senior level
150K-180K Annually
Senior level
Hardware • Healthtech • Machine Learning • Software
Lead reliability engineering for electromechanical systems, including testing and validation of hardware to ensure performance and durability.
Top Skills: Accelerated Life TestingFmeaHalt/HassJmpLabviewMatlabPythonSpcWeibull Analysis
Reposted 2 Days AgoSaved
Remote
San Francisco Bay Area, CA
Mid level
Mid level
Information Technology • Software • Database • Automation
Owner of on-prem reliability and escalations: reproduce and resolve L2/L3 issues across heterogeneous Kubernetes environments, build diagnostics and automation, improve CI and e2e test stability, establish performance baselines, harden install/upgrade flows, and write tooling in Python/Go/Rust to reduce repeat incidents.
Top Skills: BenchmarkingCiCi/CdContainersE2E TestingGoHealth ChecksHelmInstallersIntegration TestingKubernetesLoad GenerationLogsMetricsNetworkingObservabilityPackagingProfilingPythonRbacRustStorageSupport BundlesTraces
Reposted 13 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills: AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Reposted 13 Days AgoSaved
In-Office
San Francisco Bay Area, CA
350K-475K Annually
Mid level
350K-475K Annually
Mid level
Artificial Intelligence • Information Technology
Diagnose and remediate hardware, firmware, and OS issues across large GPU clusters. Own drivers, kernel interfaces, diagnostics, and firmware lifecycle. Automate reliability monitoring, analyze error rates, engage vendors and manage RMAs, and write postmortems to reduce failures and improve fleet reliability for large-scale AI experiments.
Top Skills: BmcDcgmGpuHbmIdracIpmiKubernetesLinuxLinux KernelNicNvlinkNvswitchPythonRedfishRustSlurm
Reposted 13 Days AgoSaved
In-Office
San Francisco Bay Area, CA
110K-150K Annually
Mid level
110K-150K Annually
Mid level
Energy • Renewable Energy
Drive cell-package and module-level reliability testing for perovskite–silicon tandem modules. Design and execute environmental and accelerated life tests (damp heat, thermal cycling, UV, outdoor exposure), perform failure analysis, develop accelerated stress protocols, coordinate third-party testing, and provide data-driven lifetime modeling and cross-functional reliability guidance to R&D, process, and product teams.
Top Skills: 8DDamp Heat TestingDfrFailure AnalysisFmeaHeterojunction (Hjt)JmpLifetime ModelingOutdoor Exposure TestingPerovskite-Silicon TandemPv Module Reliability TestingRcaThermal CyclingUv Testing
Reposted 9 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
200K-230K Annually
Senior level
200K-230K Annually
Senior level
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills: Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
10 Days AgoSaved
Remote
San Francisco Bay Area, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 10 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Reposted 20 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
167K-226K Annually
Senior level
167K-226K Annually
Senior level
Security • Software • Cybersecurity • Automation
As a Senior Site Reliability Engineer, you will enhance the reliability of Drata’s product teams through automation, architecture reviews, and operational excellence using cloud-native technologies.
Top Skills: AiopsAWSBashDatadogDockerGitGithub ActionsKubernetesLinuxMySQLPythonTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account