Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in San Francisco, CA
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills:
AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform
Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
Build and operate reliable cloud infrastructure, developer platforms, CI/CD systems, observability, and production workloads. Support applications, data systems, ML pipelines, and AI workloads across development, staging, and production. Establish SLOs, monitoring, incident response, automation, infrastructure-as-code practices, and operational standards. Collaborate with Product Engineering, Data Engineering, Data Science, and Security while mentoring engineers and participating in support rotations.
Top Skills:
AWSAzureCi/CdDockerGCPGitInfrastructure As CodeKubernetesOpentofuPythonSnowflakeTerraformTerragruntVercelVirtual Networking
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads the design and development of AI-driven enterprise and cloud security solutions. Manages teams delivering data, analytics, and machine learning engineering projects; oversees AI model implementation, data infrastructure, pipelines, and platform deployments. Uses Python and C++ for algorithm development, analyzes data, improves data quality, ensures technical compliance, mentors staff, manages client expectations, and drives innovation across cybersecurity and privacy initiatives.
Top Skills:
AWSC++DatabricksGCPAzurePythonSnowflake
Reposted 7 Days AgoSaved
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills:
Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Design and operate Kubernetes and GCP systems, establish SRE practices including SLIs, SLOs, and error budgets, manage incidents and vulnerabilities, strengthen cloud security, and mentor engineers across teams.
Top Skills:
Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonSecrets ManagementSecure Software Development LifecycleShellTerraform
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills:
AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills:
AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
Fintech • Real Estate • PropTech
Lead technical strategy and implementation for Redfin's production database and storage systems. Architect, scale, and operate cloud database and storage services (self-managed and AWS managed), drive reliability, observability, security, backups, upgrades, migrations, HA and disaster recovery, lead incidents and root cause analysis, mentor engineers, evangelize AI code generation tools, and participate in on-call rotation.
Top Skills:
Anthropic Claude CodeAWSAws Aurora/RdsAws S3CursorDynamoDBElasticacheGithub CopilotLinuxOpensearchPostgresPython
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Operate and improve reliability of large-scale database infrastructure: own 24x7 production support and on-call, build reusable database IaC and automation, standardize configurations, manage cluster operations and migrations, implement monitoring/observability, and document incident response and operational playbooks.
Top Skills:
AuroraAWSCi/CdCloudFormationCloudwatchDatabase Performance InsightsDatadogDevops GuruDynamoDBMariadbNoSQLPulumiRdsSQLTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead the establishment and maturation of the SRE function across cloud infrastructure and platform services. Define reliability targets, build observability and automation, lead incident response and root-cause analysis, improve resilience and recovery, and reduce operational toil. Partner with product and platform teams on reliability-focused design, establish incident practices, manage the SRE roadmap, and mentor engineers across Cloud Engineering and Reliability.
Top Skills:
AWSCloud-Native ObservabilityGoInfrastructure As CodeKubernetesPython
Artificial Intelligence • Information Technology
Ensure reliability of large-scale GPU supercomputing clusters by diagnosing hardware, firmware, driver, kernel, and operating-system issues. Investigate root causes, automate fleet monitoring, analyze reliability data, manage firmware qualification and rollouts, coordinate directly with hardware vendors, oversee RMAs, and improve GPU health monitoring. The role also requires writing postmortems and vendor cases and owning reliability initiatives across teams.
Top Skills:
BmcDcgmFabric ManagerGpuIdracIpmiKubernetesLinuxLinux KernelNvlinkNvswitchPythonRedfishRustSlurm
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve large-scale HPC GPU clusters for AI model training and inference. Build observability, automation, CI/CD, and incident response tooling; lead on-call rotations, ensure security/compliance, collaborate with ML engineers, and optimize capacity and costs for production ML workloads.
Top Skills:
AWSAzureBashCi/CdContainer OrchestrationDatadogDockerGCPGoGpu ClustersGrafanaHpcInfrastructure-As-CodeKubernetesKubernetes OperatorsOpentelemetryPython
Security • Cybersecurity
Plans, schedules, and coordinates facility maintenance work, including labor, parts, tools, permits, and job packages. Manages CMMS work orders, backlog, spare parts, preventive and predictive maintenance, and critical equipment reliability initiatives. Performs root cause analysis, develops procedures, tracks maintenance metrics, supports safety compliance, and coordinates resources with maintenance and operations teams to improve productivity, equipment availability, and mean time between failures.
Top Skills:
5SCmmsEpaLean ManufacturingMaximoExcelMicrosoft PowerpointMicrosoft ProjectMicrosoft WordOshaSAPSix SigmaTotal Productive MaintenanceToyota Production System
13 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Healthtech • Social Impact • Software
Define and scale reliability practices across the company by creating SLO/SLA frameworks, improving observability, evolving incident response, building self-service tooling and scorecards, and driving cross-team adoption to enable teams to build and operate reliable production systems at scale.
Top Skills:
AWSDatadogEksKubernetesPostgresTerraform
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills:
Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
Fintech • Real Estate • PropTech
Lead architecture and implementation of Redfin's cloud database and storage systems with focus on reliability, scalability, observability, security, and disaster recovery. Own incident response, root cause analysis, upgrades, backups, migrations, and mentor engineers; participate in on-call rotation and collaborate with cross-functional senior leaders.
Top Skills:
Anthropic Claude CodeAWSAws AuroraAws RdsAws S3CursorDynamoDBElasticacheGithub CopilotLinuxOpensearchPostgresPython
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills:
Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Information Technology
Owns reliability analysis for mission-critical electrical systems in data centers. Develops electrical BIA and FMEA packages, defines monitoring and maintenance strategies, validates one-line diagrams and redundancy assumptions, and recommends testing, spare-parts, redesign, and risk-reduction actions. Supports root-cause analyses, commissioning, vendor reviews, workshops, and design changes while communicating technical risks and tracking reliability improvements. Requires travel of 25% to 40% and occasional maintenance-window support.
Top Skills:
BatteriesCondition MonitoringElectrical DistributionFmeaGeneratorsNeta TestingOne-Line DiagramsPower SystemsProtection And ControlsSwitchgearTransfer SystemsUps Systems
Energy • Manufacturing • Solar • Renewable Energy
Leads reliability engineering for power generation, wind, electrification, data center, hybrid, and energy storage systems. Develops reliability models and simulations, coordinates equipment and protective relay requirements, evaluates architectures, and translates customer availability goals into engineering designs. Supports energy policy analysis, customer reports, proposals, reference architectures, and cross-functional projects involving engineering, sales, government, universities, and industry partners.
Top Skills:
BessBlocksimFmecaFtaHv/Lv Electrical GearIeee/Ansi StandardsMapsMarsPlcProtective RelayingRam AnalysisRcaScadaStatcom
Fintech
Provides enterprise-level technical leadership to improve the reliability, resilience, scalability, observability, and operational health of production services. The role drives automation, DevOps, CI/CD, infrastructure as code, incident response, capacity management, resilience testing, and service-health practices. Partners across engineering teams to address systemic production risks, establish technical direction, reduce operational toil, and improve sustainable on-call and recovery processes.
Top Skills:
AlertingAWSCi/CdDistributed SystemsDockerError BudgetsGoogle Cloud PlatformInfrastructure As CodeKubernetesLinuxAzureMonitoringObservabilityOracle Cloud InfrastructureSlisSlosUnix
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results

















.png)














