Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in San Francisco, CA
Aerospace • Hardware • Logistics • Robotics • Software • Transportation
Own reliability strategy for mission-critical electronic subsystems, apply DfR and PoF, lead DFMEA, design and run ALT and environmental tests, perform root-cause failure analysis, define life limits and service intervals, and improve reliability processes across hardware and operations.
Top Skills:
Accelerated Life Testing (Alt)Analog CircuitsDesign-For-Reliability (Dfr)DfmeaDigital CircuitsElectromechanical SystemsPcb LayoutPhysics-Of-Failure (Pof)Power ElectronicsRfSchematic Design
Software
Lead ownership and support of production database systems (MySQL CloudSQL and Spanner). Drive schema change processes, build automated deploy/test pipelines, implement monitoring/alerting, manage backup/recovery and DR, perform production debugging/RCA, and evaluate database features for global scalability and uptime.
Top Skills:
Cloud SqlGoogle Cloud Platform (Gcp)Google Cloud SpannerInfrastructure As CodeMysql 8.4RedisShardingSQL
Robotics
Drive hardware reliability for autonomous construction machines: define reliability strategy and requirements, lead DFMEAs and FTAs, build reliability models and tests (HALT/HASS, ALT), run FRACAS and root-cause investigations, qualify suppliers, and convert field failures into design and process improvements.
Top Skills:
MinitabPdmPlmPythonRRelexReliasoft BlocksimReliasoft Weibull++Windchill Quality
Aerospace • Hardware • Logistics • Robotics • Software • Transportation
Define and own reliability strategy for mechanical subsystems, apply DfR/PoF, lead DFMEA, design accelerated and environmental tests, run failure analysis and life modeling, set verification criteria, and drive cross-functional reliability decisions to ensure hardware readiness for flight.
Top Skills:
Accelerated Life Testing (Alt)Design-For-Reliability (Dfr)DfmeaElectromechanical SystemsEnvironmental TestingFailure AnalysisLife ModelingPhysics-Of-Failure (Pof)
Software
Lead reliability activities for photonic integrated circuits (PICs): evaluate failure modes, coordinate accelerated stress tests, develop life models from aging-data, and drive failure mode analyses across design, development, and production teams.
Information Technology • Consulting
Provide platform reliability and operational support for AEM as a Cloud Service and Vercel hosting platforms. Monitor health and performance, troubleshoot incidents, manage CI/CD pipelines and deployments, maintain integrations (DAM, Workfront, Azure, Akamai), enforce governance/security, perform capacity planning, run on-call incident response, and produce operational runbooks and RCAs. Drive platform maturity, automation, and stability for modern web applications (Next.js/React).
Top Skills:
AemAem As A Cloud ServiceAkamaiAPIsAzureCdnCi/CdCloud DeploymentsDamDispatcherMonitoring/ObservabilityNext.JsPlatform AutomationReactVercelWorkfront
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Lead design and delivery of scalable cloud infrastructure for the Spend product. Embed with development teams to drive reliability, performance, observability, incident response, and automation. Own SLOs, runbooks, DevOps metrics, and collaborate with central DevOps and security teams to ensure compliance and resilience. Lead infrastructure projects including new service launches, data centre migrations, and modernising data pipelines.
Top Skills:
Analytics PipelinesAWSData StreamingDevOpsGCPIncident ResponseKubernetesObservabilitySlosSre
Artificial Intelligence • Software
Lead reliability engineering for a reference design: build and validate availability and RAM models, run cross-discipline FMEAs, quantify failure rates and redundancy trade-offs, and close the loop by feeding fleet field-failure data back into designs to improve maintainability and uptime.
Top Skills:
Availability ModelingData Center TopologiesFmeaRam Modeling SoftwareWeibull Analysis
Software
As a Senior DevOps / Platform Reliability Engineer, you will manage CI/CD pipelines, automate infrastructure, operate Kubernetes, and enhance observability while ensuring security and compliance for enterprise systems.
Top Skills:
Argo CdAurora MysqlAWSBashCloudFormationEksElasticacheGithub ActionsGrafanaKubernetesLinuxMskOpentelemetryPrometheusPythonS3Terraform
Renewable Energy
Own reliability, performance, and scalability of Postgres and ClickHouse databases. Build scalable data pipelines, design analytical schemas and DBT models, migrate data to ClickHouse, implement data quality checks, eliminate duplicates, and manage database infrastructure via IaC.
Top Skills:
Aws CdkAws Step FunctionsCi/CdClickhouseDagsterDbtPostgresPulumiPythonSQL
Artificial Intelligence • Software
Own reliability for named customer workloads; debug distributed systems across hardware, fabric, and scheduler; run customer-facing incident communications; convert recurring customer pain into engineering fixes; support large-scale compute customers and push internal teams to resolve root causes.
Top Skills:
CloudGpuHpcInfinibandKubernetesNcclRoceSlurm
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Insurance • Software • Automation
Lead design, automation, and optimization of database infrastructure (PostgreSQL/Aurora). Build monitoring, tuning, and scaling strategies, create automation tooling, drive performance and reliability initiatives, and expand into broader SRE responsibilities to improve availability and system health for a growing SaaS platform.
Top Skills:
Amazon AuroraCi/CdDockerJavaScriptKubernetesNode.jsPostgresPrismaRedshiftTerraformTerragruntTypescript
Automotive • Information Technology • Other • Transportation • Energy
Perform RAM and FMECA/FMEA analyses, develop fault trees and reliability predictions, support maintainability and logistics analyses, produce reliability growth test plans, contribute to systems engineering documentation, advise design engineers on R&M shortfalls, and present results to management and clients.
Top Skills:
Fault Tree AnalysisFmeaFmecaIntegrated Logistics Support (Ils/Ilsa)Iso-9000Mil-Hdbk-217FRam ModellingRam SoftwareStatistical Methods
Artificial Intelligence • Software
The Site Reliability Engineer ensures the reliability and performance of products Devin and Windsurf, managing incident response, CI/CD pipelines, infrastructure as code, and fostering a reliability culture within the engineering team.
Top Skills:
AWSAzureCi/CdGCPKubernetesTerraform
Reposted 19 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Productivity
The Senior Site Reliability Engineer will enhance site reliability through monitoring, optimizing infrastructure, collaborating on engineering projects, and ensuring systems’ stability.
Top Skills:
AWSDockerKubernetesTemporal
Artificial Intelligence • Machine Learning • Generative AI
The Software Engineer in Reliability will ensure system scalability, reliability, and performance, collaborating with teams to improve infrastructure and handle incidents.
Top Skills:
Cloud InfrastructureCloudFormationContainer Orchestration PlatformsContainerization TechnologiesDatadogGrafanaIac ToolsKubernetesMicroservices ArchitectureObservability ToolsProgramming LanguagesPrometheusService Mesh TechnologiesSplunkTerraform
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
The Site Reliability Engineer will develop, deploy, and operate AI infrastructure, focusing on high-performance and scalable machine learning systems using Kubernetes and cloud platforms.
Top Skills:
AWSAzureC++GCPGoKubernetesOci
Artificial Intelligence • Natural Language Processing • Generative AI
The role involves enhancing AI system reliability, developing service objectives, monitoring infrastructure, leading incident responses, and collaborating across teams.
Top Skills:
Ai-Specific Observability ToolsDistributed SystemsHigh-Availability InfrastructureMl Hardware AcceleratorsMonitoring And Observability Systems
Artificial Intelligence • Machine Learning • Database
The role involves ensuring the reliability and performance of distributed database systems, developing monitoring strategies, and automating operations in a cloud-native environment.
Top Skills:
AnsibleArgoAWSAzureDockerGCPGitlab CiGoJavaJenkinsKubernetesPythonTerraform
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Healthtech • Insurance
Lead design, development, and operation of cloud infrastructure and SRE-focused systems. Own medium-to-large infrastructure projects, build resilient platforms, and drive cross-team technical delivery. Mentor engineers, define SLOs, reduce failure domains, and build tooling for automated, secure CI/CD and production reliability.
Top Skills:
ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
Artificial Intelligence • Healthtech • Information Technology • Software
As a Site Reliability Engineer, you will manage the production environment, focusing on infrastructure design, automation, and optimizing deployment pipelines to ensure high availability.
Top Skills:
HelmKafkaKubernetesPostgresPythonRedisTerraformTypescript
Cloud • Digital Media • Information Technology
Operate and improve Kubernetes-based production systems, manage cluster lifecycle and networking, build CI/CD and GitOps pipelines, define SLOs and incident response, automate resolution with AI, implement monitoring/alerting, and drive reliability through automation and chaos engineering.
Top Skills:
AnsibleArgocdBashBgpCalicoCephCiliumCni PluginsCorootDatadogDnsEbpfFalcoFluxcdGoGrafanaKubernetesLokiLonghornMetallbPrometheusPythonSIEMTerraformThanosVictoriametricsVxlanXdp
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
All Filters
Total selected ()
No Results
No Results



















.png)














