Top Reliability Engineer Jobs in San Francisco, CA

Reposted 11 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
168K-210K Annually
Senior level
168K-210K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills: AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
12 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
225K-265K Annually
Senior level
225K-265K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills: AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform
3 Days AgoSaved
Easy Apply
Hybrid
San Francisco Bay Area, CA
Easy Apply
186K-232K Annually
Senior level
186K-232K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
Build and operate reliable cloud infrastructure, developer platforms, CI/CD systems, observability, and production workloads. Support applications, data systems, ML pipelines, and AI workloads across development, staging, and production. Establish SLOs, monitoring, incident response, automation, infrastructure-as-code practices, and operational standards. Collaborate with Product Engineering, Data Engineering, Data Science, and Security while mentoring engineers and participating in support rotations.
Top Skills: AWSAzureCi/CdDockerGCPGitInfrastructure As CodeKubernetesOpentofuPythonSnowflakeTerraformTerragruntVercelVirtual Networking
Reposted 7 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
99K-232K Annually
Senior level
99K-232K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads the design and development of AI-driven enterprise and cloud security solutions. Manages teams delivering data, analytics, and machine learning engineering projects; oversees AI model implementation, data infrastructure, pipelines, and platform deployments. Uses Python and C++ for algorithm development, analyzes data, improves data quality, ensures technical compliance, mentors staff, manages client expectations, and drives innovation across cybersecurity and privacy initiatives.
Top Skills: AWSC++DatabricksGCPAzurePythonSnowflake
Reposted 7 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
124K-280K Annually
Senior level
124K-280K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills: Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
12 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
233K-292K Annually
Senior level
233K-292K Annually
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Design and operate Kubernetes and GCP systems, establish SRE practices including SLIs, SLOs, and error budgets, manage incidents and vulnerabilities, strengthen cloud security, and mentor engineers across teams.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonSecrets ManagementSecure Software Development LifecycleShellTerraform
5 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
6 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
7 Days AgoSaved
Remote
San Francisco Bay Area, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills: AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
7 Days AgoSaved
Remote
San Francisco Bay Area, CA
170K-225K Annually
Mid level
170K-225K Annually
Mid level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills: AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
One Month AgoSaved
Hybrid
San Francisco Bay Area, CA
180K-279K Annually
Senior level
180K-279K Annually
Senior level
Fintech • Real Estate • PropTech
Lead technical strategy and implementation for Redfin's production database and storage systems. Architect, scale, and operate cloud database and storage services (self-managed and AWS managed), drive reliability, observability, security, backups, upgrades, migrations, HA and disaster recovery, lead incidents and root cause analysis, mentor engineers, evangelize AI code generation tools, and participate in on-call rotation.
Top Skills: Anthropic Claude CodeAWSAws Aurora/RdsAws S3CursorDynamoDBElasticacheGithub CopilotLinuxOpensearchPostgresPython
Reposted One Month AgoSaved
Hybrid
San Francisco Bay Area, CA
203K-254K Annually
Senior level
203K-254K Annually
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Operate and improve reliability of large-scale database infrastructure: own 24x7 production support and on-call, build reusable database IaC and automation, standardize configurations, manage cluster operations and migrations, implement monitoring/observability, and document incident response and operational playbooks.
Top Skills: AuroraAWSCi/CdCloudFormationCloudwatchDatabase Performance InsightsDatadogDevops GuruDynamoDBMariadbNoSQLPulumiRdsSQLTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
18 Days AgoSaved
In-Office
San Francisco Bay Area, CA
220K-340K Annually
Senior level
220K-340K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead the establishment and maturation of the SRE function across cloud infrastructure and platform services. Define reliability targets, build observability and automation, lead incident response and root-cause analysis, improve resilience and recovery, and reduce operational toil. Partner with product and platform teams on reliability-focused design, establish incident practices, manage the SRE roadmap, and mentor engineers across Cloud Engineering and Reliability.
Top Skills: AWSCloud-Native ObservabilityGoInfrastructure As CodeKubernetesPython
14 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
350K-475K Annually
Entry level
350K-475K Annually
Entry level
Artificial Intelligence • Information Technology
Ensure reliability of large-scale GPU supercomputing clusters by diagnosing hardware, firmware, driver, kernel, and operating-system issues. Investigate root causes, automate fleet monitoring, analyze reliability data, manage firmware qualification and rollouts, coordinate directly with hardware vendors, oversee RMAs, and improve GPU health monitoring. The role also requires writing postmortems and vendor cases and owning reliability initiatives across teams.
Top Skills: BmcDcgmFabric ManagerGpuIdracIpmiKubernetesLinuxLinux KernelNvlinkNvswitchPythonRedfishRustSlurm
Reposted 5 Days AgoSaved
Remote
San Francisco Bay Area, CA
120K-304K Annually
Mid level
120K-304K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve large-scale HPC GPU clusters for AI model training and inference. Build observability, automation, CI/CD, and incident response tooling; lead on-call rotations, ensure security/compliance, collaborate with ML engineers, and optimize capacity and costs for production ML workloads.
Top Skills: AWSAzureBashCi/CdContainer OrchestrationDatadogDockerGCPGoGpu ClustersGrafanaHpcInfrastructure-As-CodeKubernetesKubernetes OperatorsOpentelemetryPython
5 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
48-48 Hourly
Junior
48-48 Hourly
Junior
Security • Cybersecurity
Plans, schedules, and coordinates facility maintenance work, including labor, parts, tools, permits, and job packages. Manages CMMS work orders, backlog, spare parts, preventive and predictive maintenance, and critical equipment reliability initiatives. Performs root cause analysis, develops procedures, tracks maintenance metrics, supports safety compliance, and coordinates resources with maintenance and operations teams to improve productivity, equipment availability, and mean time between failures.
Top Skills: 5SCmmsEpaLean ManufacturingMaximoExcelMicrosoft PowerpointMicrosoft ProjectMicrosoft WordOshaSAPSix SigmaTotal Productive MaintenanceToyota Production System
13 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Reposted One Month AgoSaved
Hybrid
San Francisco Bay Area, CA
182K-250K Annually
Senior level
182K-250K Annually
Senior level
Healthtech • Social Impact • Software
Define and scale reliability practices across the company by creating SLO/SLA frameworks, improving observability, evolving incident response, building self-service tooling and scorecards, and driving cross-team adoption to enable teams to build and operate reliable production systems at scale.
Top Skills: AWSDatadogEksKubernetesPostgresTerraform
24 Days AgoSaved
Easy Apply
Hybrid
San Francisco Bay Area, CA
Easy Apply
150K-220K Annually
Expert/Leader
150K-220K Annually
Expert/Leader
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills: Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
One Month AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
180K-279K Annually
Senior level
180K-279K Annually
Senior level
Fintech • Real Estate • PropTech
Lead architecture and implementation of Redfin's cloud database and storage systems with focus on reliability, scalability, observability, security, and disaster recovery. Own incident response, root cause analysis, upgrades, backups, migrations, and mentor engineers; participate in on-call rotation and collaborate with cross-functional senior leaders.
Top Skills: Anthropic Claude CodeAWSAws AuroraAws RdsAws S3CursorDynamoDBElasticacheGithub CopilotLinuxOpensearchPostgresPython
One Month AgoSaved
Remote
San Francisco Bay Area, CA
180K-270K Annually
Senior level
180K-270K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills: Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
19 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
191K-226K Annually
Senior level
191K-226K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills: AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
16 Days AgoSaved
Remote
San Francisco Bay Area, CA
Mid level
Mid level
Information Technology
Owns reliability analysis for mission-critical electrical systems in data centers. Develops electrical BIA and FMEA packages, defines monitoring and maintenance strategies, validates one-line diagrams and redundancy assumptions, and recommends testing, spare-parts, redesign, and risk-reduction actions. Supports root-cause analyses, commissioning, vendor reviews, workshops, and design changes while communicating technical risks and tracking reliability improvements. Requires travel of 25% to 40% and occasional maintenance-window support.
Top Skills: BatteriesCondition MonitoringElectrical DistributionFmeaGeneratorsNeta TestingOne-Line DiagramsPower SystemsProtection And ControlsSwitchgearTransfer SystemsUps Systems
17 Days AgoSaved
Remote
San Francisco Bay Area, CA
152K-254K Annually
Senior level
152K-254K Annually
Senior level
Energy • Manufacturing • Solar • Renewable Energy
Leads reliability engineering for power generation, wind, electrification, data center, hybrid, and energy storage systems. Develops reliability models and simulations, coordinates equipment and protective relay requirements, evaluates architectures, and translates customer availability goals into engineering designs. Supports energy policy analysis, customer reports, proposals, reference architectures, and cross-functional projects involving engineering, sales, government, universities, and industry partners.
Top Skills: BessBlocksimFmecaFtaHv/Lv Electrical GearIeee/Ansi StandardsMapsMarsPlcProtective RelayingRam AnalysisRcaScadaStatcom
4 Days AgoSaved
In-Office
San Francisco Bay Area, CA
173K-276K Annually
Expert/Leader
173K-276K Annually
Expert/Leader
Fintech
Provides enterprise-level technical leadership to improve the reliability, resilience, scalability, observability, and operational health of production services. The role drives automation, DevOps, CI/CD, infrastructure as code, incident response, capacity management, resilience testing, and service-health practices. Partners across engineering teams to address systemic production risks, establish technical direction, reduce operational toil, and improve sustainable on-call and recovery processes.
Top Skills: AlertingAWSCi/CdDistributed SystemsDockerError BudgetsGoogle Cloud PlatformInfrastructure As CodeKubernetesLinuxAzureMonitoringObservabilityOracle Cloud InfrastructureSlisSlosUnix
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account