Top Reliability Engineer Jobs in San Francisco, CA

5 Days AgoSaved
In-Office
San Francisco Bay Area, CA
85K-102K Annually
Junior
85K-102K Annually
Junior
Fintech • Payments
Develop and maintain reliable platforms through automation, observability, incident response, performance optimization, and infrastructure improvements. The role collaborates with engineering teams, supports 24x7 reliability rotations, troubleshoots complex issues across code and infrastructure, and improves operational processes. Required experience includes SRE or equivalent work, software development, cloud platforms, monitoring, databases, and containerization. Preferred skills include Terraform, REST APIs, Grafana, Splunk, Agile, and GitOps.
Top Skills: AWSAzureC#DockerGCPGitopsGoGrafanaJavaKubernetesNoSQLPythonRdbmsRest ApisSplunkTerraform
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 28 Days AgoSaved
In-Office
San Francisco Bay Area, CA
160K-220K Annually
Senior level
160K-220K Annually
Senior level
Cloud
The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.
Top Skills: AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform
Reposted 29 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
150K-180K Annually
Senior level
150K-180K Annually
Senior level
Aerospace • Automation
The Senior Reliability Engineer will establish AeroVect's reliability engineering practice, leading reliability analyses and managing external testing programs to ensure product durability and performance across operational environments.
Top Skills: Accelerated Life TestingData AnalysisEnvironmental TestingFmeaFtaMtbfMttrRbd
Reposted One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
8 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
180K-230K Annually
Senior level
180K-230K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Automation
Own the reliability, scalability, and observability of production systems. Design and operate AWS infrastructure with Terraform and Kubernetes, build monitoring and SLOs, automate deployments and capacity management, ensure database and data pipeline reliability, partner on architecture reviews, and drive infrastructure security, compliance, and incident response.
Top Skills: AWSBashCi/CdCloud PlatformsDatadogDistributed SystemsGithub ActionsGrafanaInfrastructure As CodeKubernetesPrometheusPythonTerraform
One Month AgoSaved
In-Office
San Francisco Bay Area, CA
160K-200K Annually
Mid level
160K-200K Annually
Mid level
Robotics
Own the reliability of a deployed humanoid robot fleet by diagnosing hardware, firmware, controls, and telemetry failures. Perform hands-on actuator and electronics repair, design test fixtures and mechanical modifications, build fleet health metrics and calibration tooling, analyze robot logs, improve observability, validate firmware and software changes, and document procedures. The role requires strong Python, motor-control and actuator knowledge, hardware troubleshooting, time-series analysis, and cross-functional collaboration.
Top Skills: CanDatadogGitMcapPython
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
180K-230K Annually
Senior level
180K-230K Annually
Senior level
Energy • Renewable Energy
The Staff Reliability Engineer will ensure hardware reliability in high-voltage electronics, develop reliability test programs, and collaborate on design and testing across teams.
Top Skills: Hv ElectronicsPower ConversionPython
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
150K-180K Annually
Senior level
150K-180K Annually
Senior level
Hardware • Healthtech • Machine Learning • Software
Lead reliability engineering for electromechanical systems, including testing and validation of hardware to ensure performance and durability.
Top Skills: Accelerated Life TestingFmeaHalt/HassJmpLabviewMatlabPythonSpcWeibull Analysis
Reposted One Month AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
Reposted YesterdaySaved
Remote
San Francisco Bay Area, CA
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
YesterdaySaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Edtech • Fintech • Information Technology • Software
Operate and improve AWS production infrastructure, infrastructure as code, Kubernetes workloads, observability, CI/CD, and incident response. The role investigates root causes, strengthens application resilience, automates operational tasks, supports database and performance reliability, maintains documentation, and participates in 24/7 on-call rotations. The engineer owns scoped reliability projects and collaborates with product engineering teams on resilient, secure, and compliant systems.
Top Skills: Amazon EksAmazon RdsAWSCircleCIDatadogGithub ActionsKubernetesLinuxNew RelicOpensearchPostgresRedisRubyRuby On RailsTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
YesterdaySaved
Remote
San Francisco Bay Area, CA
120K-130K Annually
Senior level
120K-130K Annually
Senior level
Hardware • Healthtech
Owns the reliability, security, performance, and availability of AWS-hosted healthcare infrastructure. Builds infrastructure as code, CI/CD automation, monitoring, observability, backup and disaster recovery capabilities. Leads incident response, optimizes cloud resources, implements security controls, supports customer onboarding and migrations, and ensures compliance with healthcare privacy and software lifecycle requirements. Participates in on-call rotations and provides technical guidance to engineering and support teams.
Top Skills: AWSBashCi/CdCitrixEcsGdprHipaaHyper-VIec 62304JavaScriptJinjaJSONMirth ConnectPythonTerraformTypescriptVMwareYaml
Reposted YesterdaySaved
Remote
San Francisco Bay Area, CA
134K-184K Annually
Senior level
134K-184K Annually
Senior level
Healthtech
Lead the migration from legacy Azure services to a Kubernetes-based, containerized microservices platform. Design, build, and scale infrastructure, implement observability (monitoring/alerting/logging), drive incident response and SLOs, automate with IaC and CI/CD, optimize cost and networking, mentor teams, and document systems to ensure reliable, scalable healthcare platform operations.
Top Skills: .NetAWSAzureAzure Entra IdBashC#DatadogGCPGithub ActionsGitlab Ci/CdGrafanaHelmKubernetesPrometheusPythonTerraform
10 Days AgoSaved
In-Office
San Francisco Bay Area, CA
150K-225K Annually
Mid level
150K-225K Annually
Mid level
Gaming
Develop and operate reliable, scalable data platforms across AWS and GCP. Build Go-based infrastructure automation, Infrastructure as Code, observability, deployment, scaling, failover, backup, and recovery tooling for Cassandra, Aerospike, Kafka, and Redis. Troubleshoot distributed systems, participate in on-call and incident response, improve service reliability, and collaborate across engineering and platform teams.
Top Skills: AerospikeAlertingAmazon ElasticacheAmazon MskAnsibleAWSCassandraConfiguration ManagementDashboardsDistributed SystemsDynamoDBGCPGoGoogle MemorystoreInfrastructure As CodeKafkaKubernetesLinuxLoggingMetricsObservabilityRedisTerraformTracing
YesterdaySaved
Remote
San Francisco Bay Area, CA
120K-140K Annually
Mid level
120K-140K Annually
Mid level
Information Technology • Consulting
Administer and secure the organization’s GitHub environment, including repositories, permissions, branch protections, security controls, and CI/CD workflows. Build automation with GitHub Actions, APIs, and scripting; monitor reliability against SLOs; troubleshoot incidents; and improve developer experience. Integrate identity providers and security tools, support migrations, maintain documentation, and guide teams on GitHub usage and Copilot adoption. The role requires SRE or DevOps experience, infrastructure-as-code, containers, cloud platforms, and observability tooling.
Top Skills: AnsibleAWSAzureBashCodeqlDatadogDependabotDockerGCPGithub ActionsGithub ApiGithub CliGithub CopilotGithub EnterpriseGrafanaPowershellPrometheusPythonSAMLScimSplunkSsoTerraform
YesterdaySaved
Remote
San Francisco Bay Area, CA
156K-199K Annually
Senior level
156K-199K Annually
Senior level
Cybersecurity
Owns the FedRAMP-authorized AWS GovCloud environment, ensuring reliability, security, compliance, patching, vulnerability remediation, continuous monitoring, deployments, hardening, audits, incident response, and on-call support. The role requires hands-on Kubernetes, AWS, infrastructure-as-code, CI/CD, vulnerability management, automation, and regulated-environment experience, while collaborating with SRE and security teams.
Top Skills: AnsibleAws GovcloudCloudFormationFips-Validated CryptographyGithub ActionsGoJenkinsKubernetesNessusOpenscapPythonQualysRubySIEMTenableTerraform
One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills: Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
Reposted 11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
131K-192K Annually
Senior level
131K-192K Annually
Senior level
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills: AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Reposted One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills: AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
194K-284K Annually
Expert/Leader
194K-284K Annually
Expert/Leader
Fintech
Provides enterprise-level technical leadership for site reliability engineering across production services. Improves reliability, scalability, observability, deployment, automation, incident response, resilience, capacity management, and operational readiness. Establishes SLIs, SLOs, error budgets, and technical direction while partnering with software engineering teams. Leads critical incident response, reduces operational toil, and develops reusable engineering practices, tooling, and platforms that improve organizational capability.
Top Skills: AlertingAWSCi/CdContainersDistributed SystemsGoogle Cloud PlatformInfrastructure As CodeLinuxAzureMonitoringObservabilityOracle Cloud InfrastructureOrchestrationScriptingSoftware AutomationUnix
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Expert/Leader
127K-249K Annually
Expert/Leader
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills: AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
Reposted One Month AgoSaved
Remote
San Francisco Bay Area, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
106K-156K Annually
Senior level
106K-156K Annually
Senior level
Fintech
The Senior Site Reliability Engineer improves the reliability, resilience, scalability, observability, and operational health of production services. Responsibilities include defining SLIs and SLOs, enhancing monitoring and incident response, automating infrastructure and deployments, reducing operational toil, troubleshooting distributed systems, and partnering with software engineering teams. The role owns reliability outcomes across multiple services, influences engineering practices, mentors colleagues, and drives continuous improvement across cloud platforms, CI/CD, infrastructure, and operational readiness.
Top Skills: AlertingAWSCi/CdContainersDistributed SystemsGoogle Cloud PlatformInfrastructure As CodeLinuxAzureMonitoringNetworkingObservabilityOracle Cloud InfrastructureOrchestrationSoftware AutomationUnix
Reposted 13 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Productivity
Build and operate reliable, scalable infrastructure and backend systems. Define and monitor SLOs and SLAs, manage on-call operations, improve observability, performance, security, developer experience, and technical processes. Contribute to Temporal-based workflow orchestration, advise backend projects, manage AWS infrastructure, and support system scalability as the customer base grows.
Top Skills: Amazon EksAWSCi/CdDockerHelmInfrastructure As CodeKubernetesLinuxTemporal
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account