Top Reliability Engineer Jobs in San Francisco, CA

One Month AgoSaved
In-Office
San Francisco Bay Area, CA
172K-203K Annually
Senior level
172K-203K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Marketing Tech • Software • Biotech • Design
Lead planning and execution of reliability testing for wearable hardware. Develop and improve test methods, coordinate with contract manufacturers and labs, perform failure analysis and DFMEA, correlate lab tests with field data, produce test plans and reports, and partner with cross-functional teams to drive corrective actions and reliability improvements.
Top Skills: DfmeaElectrical Reliability TestingEnvironmental TestingFailure AnalysisFracasMechanical Reliability TestingPfmeaReliability Test Methods
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
4 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Software
Operate and improve AWS GovCloud infrastructure for highly available, secure government systems. Responsibilities include SRE practices, incident response, Kubernetes and EKS administration, GitOps deployments with Helm and ArgoCD, Terraform automation, monitoring, security controls, compliance with FedRAMP and federal standards, change management, documentation, and cross-functional troubleshooting.
Top Skills: Amazon EksAmazon RdsAmazon S3AnsibleArgocdAws GovcloudBashDockerDocumentdbFedrampFismaGitopsGoHelmInfrastructure As CodeIso 27001KafkaKubernetesLinuxNist 800-53OpensearchPostgresPythonSoc 2Terraform
4 Days AgoSaved
Remote
San Francisco Bay Area, CA
152K-205K Annually
Senior level
152K-205K Annually
Senior level
Information Technology • Security • Software • Cybersecurity
Owns production reliability for high-throughput, low-latency systems, including observability, SLOs, incident response, capacity planning, infrastructure as code, progressive delivery, chaos testing, and operational tooling. The role requires hands-on software and infrastructure engineering, production Redis/ElastiCache expertise, cloud infrastructure knowledge, security awareness, mentoring, and participation in on-call operations.
Top Skills: AWSDatadogEksElasticacheGoGrafanaKubernetesOpentelemetryPrometheusPythonRedisTerraform
27 Days AgoSaved
Remote
San Francisco Bay Area, CA
220K-240K Annually
Senior level
220K-240K Annually
Senior level
Software
Own the reliability, performance, availability, scalability, and cost efficiency of production Aurora MySQL databases on AWS. Responsibilities include observability, alerting, query optimization, replication, backups, disaster recovery, schema migrations, reader topology, cost optimization, runbooks, incident response, and database performance reviews. The role partners with engineering teams, supports HIPAA-compliant handling of PHI, and establishes reliable database operating practices.
Top Skills: Amazon RdsAmazon RedshiftAurora MysqlAWSBashDatabricksDatadogGrafanaMySQLPercona Monitoring And Management (Pmm)Percona ToolkitPerformance InsightsPHPPrometheusPythonSnowflakeTerraform
4 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Information Technology • Software • Consulting
Build and operate reliable, observable backend systems across AWS and Python services. Responsibilities include defining SLOs and error budgets, designing monitoring and alerting, managing incident response and 24/7 on-call rotations, conducting postmortems, improving performance and capacity, automating toil reduction, maintaining infrastructure as code and deployment pipelines, and mentoring SRE engineers. The role also involves consulting with client teams, documenting operational practices, and supporting AI workloads.
Top Skills: AlbAWSBashCi/CdCloudFormationDockerEcs/FargateGitGoIamKubernetesLambdaLinux/UnixPulumiPythonRds AuroraTerraform
14 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
45-50 Hourly
Internship
45-50 Hourly
Internship
Analytics • Consulting
Supports the reliability and performance of Gallup’s global technology platform through observability, dashboards, alerting, automation, and incident response. The intern will use tools including AWS, Dynatrace, PagerDuty, Slack, and infrastructure automation technologies while collaborating with site reliability, DevOps, and software engineering teams. The position is a full-time, on-site summer internship with potential part-time school-year work.
Top Skills: Amazon EcsAWSBashDockerDynatraceGitGrafanaLinuxPagerdutyPowershellPythonSlackTerraformWindows
4 Days AgoSaved
Remote
San Francisco Bay Area, CA
250K-325K Annually
Senior level
250K-325K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Lead reliability engineering for Replit’s large-scale infrastructure by designing observability, defining SLOs and SLIs, leading incident response, automating operations, optimizing Kubernetes and GCP deployments, debugging distributed systems, and mentoring engineers. Build internal tools and integrations in Python or Go, maintain infrastructure as code and CI/CD pipelines, improve system performance and resilience, and establish reliability, security, and operational best practices across the engineering organization.
Top Skills: Ci/CdDatadogDockerGoGoogle Cloud Platform (Gcp)GrafanaKubernetesOpentelemetryPrometheusPulumiPythonTerraform
14 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
120K-249K Annually
Mid level
120K-249K Annually
Mid level
Software
Improve platform reliability, observability, resilience, and operational readiness. Partner with development teams to establish reliability standards, implement observability practices, enable autonomous service deployment and support, mentor developers, and drive adoption of incident response, SLO, error budget, and service health practices across the organization.
Top Skills: Cloud-Native PlatformsInfrastructure As CodeNode.jsObservabilityTypescript
4 Days AgoSaved
Remote
San Francisco Bay Area, CA
210K-275K Annually
Senior level
210K-275K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Design and maintain reliable, scalable infrastructure for Replit’s global platform. Responsibilities include building observability and alerting systems, automating infrastructure with infrastructure-as-code, managing CI/CD pipelines, defining SLOs and SLIs, leading incident response and postmortems, maintaining runbooks, optimizing performance, and improving capacity, availability, and recovery times.
Top Skills: AnsibleCi/CdDatadogGoGoogle Cloud PlatformGrafanaKubernetesPrometheusPulumiPythonTerraform
5 Days AgoSaved
Remote
San Francisco Bay Area, CA
101K-161K Annually
Senior level
101K-161K Annually
Senior level
Cloud • Software • Analytics
Develop and operate Arista’s FedRAMP CloudVision SaaS platform at scale. Responsibilities include improving reliability, scalability, observability, autoscaling, disaster recovery, capacity planning, CI/CD, network architecture, cost optimization, and cloud application security. The role develops and manages Kubernetes-native services, distributed databases, and automation using technologies such as GCP, GKE, Go, Python, Ansible, Pulumi, and Bash. Participation in a FedRAMP on-call rotation is required.
Top Skills: AnsibleBashCi/CdDistributed DatabasesFedrampGoGoogle Cloud Platform (Gcp)Google Kubernetes Engine (Gke)KubernetesPulumiPythonSaaS
5 Days AgoSaved
Remote
San Francisco Bay Area, CA
85K-141K Annually
Senior level
85K-141K Annually
Senior level
Cloud • Security • Cybersecurity
Owns operational capabilities for FedRAMP-regulated cloud environments, including monitoring, alerting, backup and recovery, continuous compliance evidence, incident response, and automation. The role leads major incident command, improves runbooks and operational standards, supports client-facing service delivery, partners on transitions to managed operations, and mentors engineers. Requires deep cloud operations, observability, infrastructure-as-code, resilience engineering, security-control frameworks, and senior escalation experience.
Top Skills: AnsibleAWSAzureCi/CdGCPGoInfrastructure As CodeJSONOscalPolicy As CodePythonSIEMTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 5 Days AgoSaved
Remote
San Francisco Bay Area, CA
120K-165K Annually
Senior level
120K-165K Annually
Senior level
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills: AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Reposted One Month AgoSaved
Hybrid
San Francisco Bay Area, CA
Expert/Leader
Expert/Leader
Robotics
Drive hardware reliability for autonomous construction machines: define reliability strategy and requirements, lead DFMEAs and FTAs, build reliability models and tests (HALT/HASS, ALT), run FRACAS and root-cause investigations, qualify suppliers, and convert field failures into design and process improvements.
Top Skills: MinitabPdmPlmPythonRRelexReliasoft BlocksimReliasoft Weibull++Windchill Quality
Reposted One Month AgoSaved
In-Office
San Francisco Bay Area, CA
Senior level
Senior level
Aerospace • Hardware • Logistics • Robotics • Software • Transportation
Define and own reliability strategy for mechanical subsystems, apply DfR/PoF, lead DFMEA, design accelerated and environmental tests, run failure analysis and life modeling, set verification criteria, and drive cross-functional reliability decisions to ensure hardware readiness for flight.
Top Skills: Accelerated Life Testing (Alt)Design-For-Reliability (Dfr)DfmeaElectromechanical SystemsEnvironmental TestingFailure AnalysisLife ModelingPhysics-Of-Failure (Pof)
5 Days AgoSaved
Remote
San Francisco Bay Area, CA
140K-170K Annually
Senior level
140K-170K Annually
Senior level
eCommerce • Manufacturing
Manage Azure infrastructure and AKS clusters, build GitHub Actions CI/CD pipelines, and improve Grafana-based observability and incident response. Define SLIs, SLOs, and error budgets; maintain infrastructure as code with Pulumi; troubleshoot reliability issues; perform capacity planning and performance tuning; participate in on-call support; and document operational procedures. Collaborate with development teams to deliver scalable, reliable production systems.
Top Skills: Azure Kubernetes Service (Aks)BashDnsGithub ActionsGrafanaIstioKubernetesLoad BalancingLokiAzurePrometheusPulumiPython
15 Days AgoSaved
In-Office
San Francisco Bay Area, CA
350K-475K Annually
Mid level
350K-475K Annually
Mid level
Artificial Intelligence • Information Technology
Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning systems. Debug distributed failures across accelerators, networking, storage, schedulers, and training frameworks; build monitoring, alerting, recovery, checkpointing, and scheduling tools; improve cluster utilization and fault tolerance; support production model runs through on-call rotations and postmortems; and partner closely with research teams during active training runs.
Top Skills: C++DpoGoGpuInfinibandKubernetesLinuxNcclPpoPythonPyTorchRayRdmaRlhfSlurmTpu
15 Days AgoSaved
In-Office
San Francisco Bay Area, CA
350K-475K Annually
Entry level
350K-475K Annually
Entry level
Artificial Intelligence • Information Technology
Owns end-to-end reliability for Tinker, including CI/CD, observability, service-level objectives, incident response, multi-tenant isolation, resource scheduling, and vulnerability remediation. The role designs monitoring for distributed training systems, improves recovery and checkpointing, and operates Kubernetes clusters supporting heterogeneous GPU workloads. It collaborates closely with platform, security, engineering, and research teams to improve production resilience and utilization.
Top Skills: Ci/CdCloud InfrastructureDistributed SystemsGpu WorkloadsKubernetesLoraProduction Observability
6 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Architect, develop, test, and distribute software, services, and infrastructure supporting Akamai’s cloud hypervisor platforms. Improve observability, automate infrastructure processes, troubleshoot complex distributed-system issues, mentor engineers, and participate in on-call service restoration. The role requires deep Linux, kernel, virtualization, ARM hardware, large-scale infrastructure, DevOps, and configuration-management expertise.
Top Skills: AnsibleArmDevOpsDistributed SystemsKvm/QemuLinuxLinux KernelNested VirtualizationNvidia GraceObservability InfrastructureSaltstack
6 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
169K-305K Annually
Expert/Leader
169K-305K Annually
Expert/Leader
Cloud • Security • Software • Cybersecurity
Architect, build, and support reliable network infrastructure and automation for Akamai’s distributed cloud platform. Develop Bash and Python tooling, establish deployment standards, define SLOs, mentor engineers, and troubleshoot complex network issues. The role requires expertise in large-scale distributed systems, TCP/IP, BGP, Linux networking, configuration management, CI/CD, and open-source networking software, with participation in on-call rotations.
Top Skills: AnsibleArgocdBashBgpBirdChefFirewallsFrrGithub ActionsGoGobgpJenkinsLinux NetworkingLoad BalancingPuppetPythonRustSaltstackSlack BotsTcp/Ip
6 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Software
Own reliability, performance, scalability, and operational standards across on-premises, private-cloud, and AWS environments. Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling; define SLOs, SLIs, error budgets, and readiness standards. Lead incident response, root-cause analysis, disaster-recovery readiness, release coordination, and preventive automation. Mentor engineers, coach teams on operational practices, coordinate on-call coverage, and ensure infrastructure meets security and compliance requirements.
Top Skills: AnsibleAuto ScalingAWSBashCi/CdClaudeCloudwatchCrowdstrikeDatadogDistributed SystemsDnsDockerEc2EcsGitGithub CopilotGitlabGitopsIamKubernetesLinuxLoad BalancingOracle LinuxPythonQualysRapid7RhelS3Tcp/IpTerraformVpcWireshark
Reposted 6 Days AgoSaved
Remote
San Francisco Bay Area, CA
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills: AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Reposted 7 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills: ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
7 Days AgoSaved
Remote
San Francisco Bay Area, CA
Entry level
Entry level
Artificial Intelligence • Hardware • Software • Semiconductor
Develop kernel-centric reliability solutions for AI compute clusters and production services. Responsibilities include debugging failures, building diagnostic tools, supporting incident response, performing root-cause analysis, improving kernel and software reliability, and collaborating with systems, hardware, ASIC, and architecture teams on reliability-focused designs.
Top Skills: CC++Core Dump HandlingDebuggersDistributed ProgrammingGpusParallel ProgrammingProfilersPythonSanitizersTracing
7 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
168K-334K Annually
Senior level
168K-334K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve large-scale Kubernetes and GPU clusters across public and private clouds. Build automation, observability, capacity-management, and reliability systems; define SLOs and SLIs; support production launches; lead incident triage and root-cause analysis; conduct blameless postmortems; and participate in on-call support for AI workloads.
Top Skills: AirflowAnsibleArgo WorkflowsAWSAws Step FunctionsCadenceChefCudaElk StackGoGoogle Cloud PlatformGrafanaKubernetesKubevirtLightstepLinuxAzureNcclNvidia Dgx CloudNvidia DynamoOpentelemetryOracle Cloud InfrastructurePrometheusPuppetPythonPyTorchSglangSplunkTcp/IpTemporalTensorrt-LlmTerraformVllm
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account