Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in San Francisco, CA
Fintech • Payments
Develop and maintain reliable platforms through automation, observability, incident response, performance optimization, and infrastructure improvements. The role collaborates with engineering teams, supports 24x7 reliability rotations, troubleshoots complex issues across code and infrastructure, and improves operational processes. Required experience includes SRE or equivalent work, software development, cloud platforms, monitoring, databases, and containerization. Preferred skills include Terraform, REST APIs, Grafana, Splunk, Agile, and GitOps.
Top Skills:
AWSAzureC#DockerGCPGitopsGoGrafanaJavaKubernetesNoSQLPythonRdbmsRest ApisSplunkTerraform
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Cloud
The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.
Top Skills:
AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform
Aerospace • Automation
The Senior Reliability Engineer will establish AeroVect's reliability engineering practice, leading reliability analyses and managing external testing programs to ensure product durability and performance across operational environments.
Top Skills:
Accelerated Life TestingData AnalysisEnvironmental TestingFmeaFtaMtbfMttrRbd
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Artificial Intelligence • Information Technology • Software • Automation
Own the reliability, scalability, and observability of production systems. Design and operate AWS infrastructure with Terraform and Kubernetes, build monitoring and SLOs, automate deployments and capacity management, ensure database and data pipeline reliability, partner on architecture reviews, and drive infrastructure security, compliance, and incident response.
Top Skills:
AWSBashCi/CdCloud PlatformsDatadogDistributed SystemsGithub ActionsGrafanaInfrastructure As CodeKubernetesPrometheusPythonTerraform
Robotics
Own the reliability of a deployed humanoid robot fleet by diagnosing hardware, firmware, controls, and telemetry failures. Perform hands-on actuator and electronics repair, design test fixtures and mechanical modifications, build fleet health metrics and calibration tooling, analyze robot logs, improve observability, validate firmware and software changes, and document procedures. The role requires strong Python, motor-control and actuator knowledge, hardware troubleshooting, time-series analysis, and cross-functional collaboration.
Top Skills:
CanDatadogGitMcapPython
Energy • Renewable Energy
The Staff Reliability Engineer will ensure hardware reliability in high-voltage electronics, develop reliability test programs, and collaborate on design and testing across teams.
Top Skills:
Hv ElectronicsPower ConversionPython
Hardware • Healthtech • Machine Learning • Software
Lead reliability engineering for electromechanical systems, including testing and validation of hardware to ensure performance and durability.
Top Skills:
Accelerated Life TestingFmeaHalt/HassJmpLabviewMatlabPythonSpcWeibull Analysis
Reposted One Month AgoSaved
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills:
ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Edtech • Fintech • Information Technology • Software
Operate and improve AWS production infrastructure, infrastructure as code, Kubernetes workloads, observability, CI/CD, and incident response. The role investigates root causes, strengthens application resilience, automates operational tasks, supports database and performance reliability, maintains documentation, and participates in 24/7 on-call rotations. The engineer owns scoped reliability projects and collaborates with product engineering teams on resilient, secure, and compliant systems.
Top Skills:
Amazon EksAmazon RdsAWSCircleCIDatadogGithub ActionsKubernetesLinuxNew RelicOpensearchPostgresRedisRubyRuby On RailsTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Hardware • Healthtech
Owns the reliability, security, performance, and availability of AWS-hosted healthcare infrastructure. Builds infrastructure as code, CI/CD automation, monitoring, observability, backup and disaster recovery capabilities. Leads incident response, optimizes cloud resources, implements security controls, supports customer onboarding and migrations, and ensures compliance with healthcare privacy and software lifecycle requirements. Participates in on-call rotations and provides technical guidance to engineering and support teams.
Top Skills:
AWSBashCi/CdCitrixEcsGdprHipaaHyper-VIec 62304JavaScriptJinjaJSONMirth ConnectPythonTerraformTypescriptVMwareYaml
Healthtech
Lead the migration from legacy Azure services to a Kubernetes-based, containerized microservices platform. Design, build, and scale infrastructure, implement observability (monitoring/alerting/logging), drive incident response and SLOs, automate with IaC and CI/CD, optimize cost and networking, mentor teams, and document systems to ensure reliable, scalable healthcare platform operations.
Top Skills:
.NetAWSAzureAzure Entra IdBashC#DatadogGCPGithub ActionsGitlab Ci/CdGrafanaHelmKubernetesPrometheusPythonTerraform
Gaming
Develop and operate reliable, scalable data platforms across AWS and GCP. Build Go-based infrastructure automation, Infrastructure as Code, observability, deployment, scaling, failover, backup, and recovery tooling for Cassandra, Aerospike, Kafka, and Redis. Troubleshoot distributed systems, participate in on-call and incident response, improve service reliability, and collaborate across engineering and platform teams.
Top Skills:
AerospikeAlertingAmazon ElasticacheAmazon MskAnsibleAWSCassandraConfiguration ManagementDashboardsDistributed SystemsDynamoDBGCPGoGoogle MemorystoreInfrastructure As CodeKafkaKubernetesLinuxLoggingMetricsObservabilityRedisTerraformTracing
Information Technology • Consulting
Administer and secure the organization’s GitHub environment, including repositories, permissions, branch protections, security controls, and CI/CD workflows. Build automation with GitHub Actions, APIs, and scripting; monitor reliability against SLOs; troubleshoot incidents; and improve developer experience. Integrate identity providers and security tools, support migrations, maintain documentation, and guide teams on GitHub usage and Copilot adoption. The role requires SRE or DevOps experience, infrastructure-as-code, containers, cloud platforms, and observability tooling.
Top Skills:
AnsibleAWSAzureBashCodeqlDatadogDependabotDockerGCPGithub ActionsGithub ApiGithub CliGithub CopilotGithub EnterpriseGrafanaPowershellPrometheusPythonSAMLScimSplunkSsoTerraform
Cybersecurity
Owns the FedRAMP-authorized AWS GovCloud environment, ensuring reliability, security, compliance, patching, vulnerability remediation, continuous monitoring, deployments, hardening, audits, incident response, and on-call support. The role requires hands-on Kubernetes, AWS, infrastructure-as-code, CI/CD, vulnerability management, automation, and regulated-environment experience, while collaborating with SRE and security teams.
Top Skills:
AnsibleAws GovcloudCloudFormationFips-Validated CryptographyGithub ActionsGoJenkinsKubernetesNessusOpenscapPythonQualysRubySIEMTenableTerraform
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills:
Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills:
AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills:
AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Fintech
Provides enterprise-level technical leadership for site reliability engineering across production services. Improves reliability, scalability, observability, deployment, automation, incident response, resilience, capacity management, and operational readiness. Establishes SLIs, SLOs, error budgets, and technical direction while partnering with software engineering teams. Leads critical incident response, reduces operational toil, and develops reusable engineering practices, tooling, and platforms that improve organizational capability.
Top Skills:
AlertingAWSCi/CdContainersDistributed SystemsGoogle Cloud PlatformInfrastructure As CodeLinuxAzureMonitoringObservabilityOracle Cloud InfrastructureOrchestrationScriptingSoftware AutomationUnix
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills:
AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
Reposted One Month AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Fintech
The Senior Site Reliability Engineer improves the reliability, resilience, scalability, observability, and operational health of production services. Responsibilities include defining SLIs and SLOs, enhancing monitoring and incident response, automating infrastructure and deployments, reducing operational toil, troubleshooting distributed systems, and partnering with software engineering teams. The role owns reliability outcomes across multiple services, influences engineering practices, mentors colleagues, and drives continuous improvement across cloud platforms, CI/CD, infrastructure, and operational readiness.
Top Skills:
AlertingAWSCi/CdContainersDistributed SystemsGoogle Cloud PlatformInfrastructure As CodeLinuxAzureMonitoringNetworkingObservabilityOracle Cloud InfrastructureOrchestrationSoftware AutomationUnix
Productivity
Build and operate reliable, scalable infrastructure and backend systems. Define and monitor SLOs and SLAs, manage on-call operations, improve observability, performance, security, developer experience, and technical processes. Contribute to Temporal-based workflow orchestration, advise backend projects, manage AWS infrastructure, and support system scalability as the customer base grows.
Top Skills:
Amazon EksAWSCi/CdDockerHelmInfrastructure As CodeKubernetesLinuxTemporal
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results











.png)















.png)



