Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in San Francisco, CA
Healthtech • Social Impact • Software
Define and scale reliability practices across the company by creating SLO/SLA frameworks, improving observability, evolving incident response, building self-service tooling and scorecards, and driving cross-team adoption to enable teams to build and operate reliable production systems at scale.
Top Skills:
AWSDatadogEksKubernetesPostgresTerraform
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
The Senior Hardware Reliability Engineer ensures product reliability through planning, testing, and collaboration across engineering and operations. Responsibilities include leading investigations, analyzing failure data, and designing reliability strategies throughout the product lifecycle.
Top Skills:
Environmental TestingFailure AnalysisFirmware EngineeringHardware ReliabilityReliability ModelingStress Testing
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Operate and improve reliability of large-scale database infrastructure: own 24x7 production support and on-call, build reusable database IaC and automation, standardize configurations, manage cluster operations and migrations, implement monitoring/observability, and document incident response and operational playbooks.
Top Skills:
AuroraAWSCi/CdCloudFormationCloudwatchDatabase Performance InsightsDatadogDevops GuruDynamoDBMariadbNoSQLPulumiRdsSQLTerraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Support and improve reliability, scalability, and performance of large-scale databases across cloud and on-prem. Build automation, Kubernetes operators, GitOps workflows, and tooling in Go or Python; implement observability, capacity planning, failover, backups, schema migrations, and self-healing systems; partner with application teams and leverage AI to improve operations and engineering productivity.
Top Skills:
AerospikeArgocdAurora MysqlClaudeCursorEksFluxcdGithub CopilotGitopsGkeGoKubernetesKubernetes OperatorsMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Reposted 21 Days AgoSaved
Fintech • Financial Services
The Staff Infrastructure Reliability Engineer leads Redfin's production database and storage systems, collaborating on strategies for reliability, scalability, and performance, while mentoring engineers and guiding complex technical discussions.
Top Skills:
AWSAws AuroraAws RdsAws S3DynamoDBElasticacheOpensearchPostgresPythonRdbms
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead design, automation, and maintenance of cloud-based database infrastructure (primarily SQL Server and MySQL). Improve reliability with monitoring, HA/DR, automation, troubleshooting, on-call support, and mentoring of junior engineers while collaborating across teams.
Top Skills:
AuroraAWSBashFailover ClusteringMySQLNew RelicOrchestratorPmmPythonRdsRubySQL ServerVividcortex
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills:
Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills:
Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Financial Services • Generative AI
As a Site Reliability Engineer, you'll design and improve critical production systems, lead incident response, and enhance observability while embedding with product teams to ensure reliability and performance at scale.
Top Skills:
AWSC++Ci/CdGoPythonRust
Reposted 5 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills:
AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Hardware • Robotics • Software
Lead hardware test and reliability efforts for autonomous drones: define and execute validation strategies from prototype to production, perform failure analysis, drive corrective actions, contribute design-for-reliability, and use telemetry and data (Python/SQL/AI) to develop prognostics and reliability specs.
Top Skills:
Ai ModelsPcbaPythonSQL
6 Days AgoSaved
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills:
ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
Artificial Intelligence • Machine Learning • Software • Generative AI
Help design, scale, and improve platform reliability: define SLOs/SLIs, run on-call and incident response, build observability, improve resilience to external dependencies, enhance CI/CD and deploy safety, optimize cost and capacity, and influence infrastructure architecture.
Top Skills:
AmplitudeAWSCloud RunContainersEcsFargateFirebaseGCPKubernetesModalNext.JsNode.jsPythonReactRedisSentryServerlessTypescriptUpstash
Wearables
The Hardware Reliability Engineer will oversee reliability testing for Skip's wearable devices, manage FMEA, perform environmental stress testing, and lead root cause analyses to ensure product performance in real-world conditions.
Top Skills:
Environmental TestingFmeaHaltHassThermal ChambersVibration TablesWeibull Analysis
Transportation
The Vehicle Reliability Engineer will diagnose issues, implement workarounds, collaborate on root cause analysis, and improve vehicle reliability through methodologies like Lean and Six Sigma.
Top Skills:
Automotive Diagnostic ToolsCan BusLinux
Reposted 7 Days AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Greentech • Energy
Lead Design for Reliability for electro-mechanical plant subsystems, define reliability requirements and validation roadmaps, design and run accelerated lifetime tests (focus on industrial control panels and outdoor power electronics), perform DFMEAs and root cause analysis, mentor peers, and champion safety across cross-functional teams.
Top Skills:
Accelerated Lifetime TestingActuatorsDesign For Reliability (Dfr)DfmeaElectro-Mechanical SystemsFansFluid SystemsIndustrial Control PanelsOutdoor Power Electronics
Greentech • Energy
Lead reliability efforts for iron-air battery packs: advise on Design for Reliability, run DFMEAs, create accelerated lifetime test plans for liquid/gas and electro-mechanical subsystems, identify root causes, and support cross-functional teams while championing safety.
Top Skills:
Accelerated Lifetime TestingDesign For Reliability (Dfr)DfmeaDuctsFansGas Handling SystemsLiquid Handling SystemsMechanical Subsystem TestingPipesPumps
Artificial Intelligence • Software • Energy • Renewable Energy
The Reliability Engineer will drive reliability for Solid-State Transformers, conduct test plans, analyze data, document risks, and support hardware deployment.
Top Skills:
Ansys SherlockMatlabPythonRReliasoft
Energy
The Reliability Engineer will define reliability requirements, conduct tests, analyze designs, and collaborate on hardware reliability within energy storage products.
Top Skills:
JmpMinitabPython
2 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Fintech • Mobile • Payments • Financial Services
Lead design and delivery of a reliability platform for production systems. Own quarterly goals, guide engineers through ambiguity, collaborate with PM/design/analytics, monitor and operate services (on-call), set quality and code standards, and mentor teammates. Integrate AI/LLM tooling to improve automation, debugging, and service health.
Top Skills:
Ai FrameworksAWSKotlinKubernetesLlmsMySQLPythonReactVue
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Artificial Intelligence • Logistics • Robotics • Software
Own reliability of distributed systems across cloud (Kubernetes), edge, and on-site deployments. Build observability, monitoring, alerting, and incident response processes. Improve deployment workflows, diagnose infra/networking/distributed issues, and partner with engineering to prevent recurrence and scale repeatable deployments in imperfect real-world environments.
Top Skills:
AWSAzureContainerized SystemsGCPGrafanaKafkaKubernetesLinuxNetworkingOpentelemetryPrometheusRtspSecure TunnelsVpnsWebrtc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results













.png)


















