Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in San Francisco, CA
Artificial Intelligence • Information Technology • Machine Learning • Marketing Tech • Software • Biotech • Design
Lead planning and execution of reliability testing for wearable hardware. Develop and improve test methods, coordinate with contract manufacturers and labs, perform failure analysis and DFMEA, correlate lab tests with field data, produce test plans and reports, and partner with cross-functional teams to drive corrective actions and reliability improvements.
Top Skills:
DfmeaElectrical Reliability TestingEnvironmental TestingFailure AnalysisFracasMechanical Reliability TestingPfmeaReliability Test Methods
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Software
Operate and improve AWS GovCloud infrastructure for highly available, secure government systems. Responsibilities include SRE practices, incident response, Kubernetes and EKS administration, GitOps deployments with Helm and ArgoCD, Terraform automation, monitoring, security controls, compliance with FedRAMP and federal standards, change management, documentation, and cross-functional troubleshooting.
Top Skills:
Amazon EksAmazon RdsAmazon S3AnsibleArgocdAws GovcloudBashDockerDocumentdbFedrampFismaGitopsGoHelmInfrastructure As CodeIso 27001KafkaKubernetesLinuxNist 800-53OpensearchPostgresPythonSoc 2Terraform
Information Technology • Security • Software • Cybersecurity
Owns production reliability for high-throughput, low-latency systems, including observability, SLOs, incident response, capacity planning, infrastructure as code, progressive delivery, chaos testing, and operational tooling. The role requires hands-on software and infrastructure engineering, production Redis/ElastiCache expertise, cloud infrastructure knowledge, security awareness, mentoring, and participation in on-call operations.
Top Skills:
AWSDatadogEksElasticacheGoGrafanaKubernetesOpentelemetryPrometheusPythonRedisTerraform
Software
Own the reliability, performance, availability, scalability, and cost efficiency of production Aurora MySQL databases on AWS. Responsibilities include observability, alerting, query optimization, replication, backups, disaster recovery, schema migrations, reader topology, cost optimization, runbooks, incident response, and database performance reviews. The role partners with engineering teams, supports HIPAA-compliant handling of PHI, and establishes reliable database operating practices.
Top Skills:
Amazon RdsAmazon RedshiftAurora MysqlAWSBashDatabricksDatadogGrafanaMySQLPercona Monitoring And Management (Pmm)Percona ToolkitPerformance InsightsPHPPrometheusPythonSnowflakeTerraform
Information Technology • Software • Consulting
Build and operate reliable, observable backend systems across AWS and Python services. Responsibilities include defining SLOs and error budgets, designing monitoring and alerting, managing incident response and 24/7 on-call rotations, conducting postmortems, improving performance and capacity, automating toil reduction, maintaining infrastructure as code and deployment pipelines, and mentoring SRE engineers. The role also involves consulting with client teams, documenting operational practices, and supporting AI workloads.
Top Skills:
AlbAWSBashCi/CdCloudFormationDockerEcs/FargateGitGoIamKubernetesLambdaLinux/UnixPulumiPythonRds AuroraTerraform
Analytics • Consulting
Supports the reliability and performance of Gallup’s global technology platform through observability, dashboards, alerting, automation, and incident response. The intern will use tools including AWS, Dynatrace, PagerDuty, Slack, and infrastructure automation technologies while collaborating with site reliability, DevOps, and software engineering teams. The position is a full-time, on-site summer internship with potential part-time school-year work.
Top Skills:
Amazon EcsAWSBashDockerDynatraceGitGrafanaLinuxPagerdutyPowershellPythonSlackTerraformWindows
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Lead reliability engineering for Replit’s large-scale infrastructure by designing observability, defining SLOs and SLIs, leading incident response, automating operations, optimizing Kubernetes and GCP deployments, debugging distributed systems, and mentoring engineers. Build internal tools and integrations in Python or Go, maintain infrastructure as code and CI/CD pipelines, improve system performance and resilience, and establish reliability, security, and operational best practices across the engineering organization.
Top Skills:
Ci/CdDatadogDockerGoGoogle Cloud Platform (Gcp)GrafanaKubernetesOpentelemetryPrometheusPulumiPythonTerraform
Software
Improve platform reliability, observability, resilience, and operational readiness. Partner with development teams to establish reliability standards, implement observability practices, enable autonomous service deployment and support, mentor developers, and drive adoption of incident response, SLO, error budget, and service health practices across the organization.
Top Skills:
Cloud-Native PlatformsInfrastructure As CodeNode.jsObservabilityTypescript
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Design and maintain reliable, scalable infrastructure for Replit’s global platform. Responsibilities include building observability and alerting systems, automating infrastructure with infrastructure-as-code, managing CI/CD pipelines, defining SLOs and SLIs, leading incident response and postmortems, maintaining runbooks, optimizing performance, and improving capacity, availability, and recovery times.
Top Skills:
AnsibleCi/CdDatadogGoGoogle Cloud PlatformGrafanaKubernetesPrometheusPulumiPythonTerraform
Cloud • Software • Analytics
Develop and operate Arista’s FedRAMP CloudVision SaaS platform at scale. Responsibilities include improving reliability, scalability, observability, autoscaling, disaster recovery, capacity planning, CI/CD, network architecture, cost optimization, and cloud application security. The role develops and manages Kubernetes-native services, distributed databases, and automation using technologies such as GCP, GKE, Go, Python, Ansible, Pulumi, and Bash. Participation in a FedRAMP on-call rotation is required.
Top Skills:
AnsibleBashCi/CdDistributed DatabasesFedrampGoGoogle Cloud Platform (Gcp)Google Kubernetes Engine (Gke)KubernetesPulumiPythonSaaS
Cloud • Security • Cybersecurity
Owns operational capabilities for FedRAMP-regulated cloud environments, including monitoring, alerting, backup and recovery, continuous compliance evidence, incident response, and automation. The role leads major incident command, improves runbooks and operational standards, supports client-facing service delivery, partners on transitions to managed operations, and mentors engineers. Requires deep cloud operations, observability, infrastructure-as-code, resilience engineering, security-control frameworks, and senior escalation experience.
Top Skills:
AnsibleAWSAzureCi/CdGCPGoInfrastructure As CodeJSONOscalPolicy As CodePythonSIEMTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills:
AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Robotics
Drive hardware reliability for autonomous construction machines: define reliability strategy and requirements, lead DFMEAs and FTAs, build reliability models and tests (HALT/HASS, ALT), run FRACAS and root-cause investigations, qualify suppliers, and convert field failures into design and process improvements.
Top Skills:
MinitabPdmPlmPythonRRelexReliasoft BlocksimReliasoft Weibull++Windchill Quality
Aerospace • Hardware • Logistics • Robotics • Software • Transportation
Define and own reliability strategy for mechanical subsystems, apply DfR/PoF, lead DFMEA, design accelerated and environmental tests, run failure analysis and life modeling, set verification criteria, and drive cross-functional reliability decisions to ensure hardware readiness for flight.
Top Skills:
Accelerated Life Testing (Alt)Design-For-Reliability (Dfr)DfmeaElectromechanical SystemsEnvironmental TestingFailure AnalysisLife ModelingPhysics-Of-Failure (Pof)
eCommerce • Manufacturing
Manage Azure infrastructure and AKS clusters, build GitHub Actions CI/CD pipelines, and improve Grafana-based observability and incident response. Define SLIs, SLOs, and error budgets; maintain infrastructure as code with Pulumi; troubleshoot reliability issues; perform capacity planning and performance tuning; participate in on-call support; and document operational procedures. Collaborate with development teams to deliver scalable, reliable production systems.
Top Skills:
Azure Kubernetes Service (Aks)BashDnsGithub ActionsGrafanaIstioKubernetesLoad BalancingLokiAzurePrometheusPulumiPython
Artificial Intelligence • Information Technology
Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning systems. Debug distributed failures across accelerators, networking, storage, schedulers, and training frameworks; build monitoring, alerting, recovery, checkpointing, and scheduling tools; improve cluster utilization and fault tolerance; support production model runs through on-call rotations and postmortems; and partner closely with research teams during active training runs.
Top Skills:
C++DpoGoGpuInfinibandKubernetesLinuxNcclPpoPythonPyTorchRayRdmaRlhfSlurmTpu
Artificial Intelligence • Information Technology
Owns end-to-end reliability for Tinker, including CI/CD, observability, service-level objectives, incident response, multi-tenant isolation, resource scheduling, and vulnerability remediation. The role designs monitoring for distributed training systems, improves recovery and checkpointing, and operates Kubernetes clusters supporting heterogeneous GPU workloads. It collaborates closely with platform, security, engineering, and research teams to improve production resilience and utilization.
Top Skills:
Ci/CdCloud InfrastructureDistributed SystemsGpu WorkloadsKubernetesLoraProduction Observability
Cloud • Security • Software • Cybersecurity
Architect, develop, test, and distribute software, services, and infrastructure supporting Akamai’s cloud hypervisor platforms. Improve observability, automate infrastructure processes, troubleshoot complex distributed-system issues, mentor engineers, and participate in on-call service restoration. The role requires deep Linux, kernel, virtualization, ARM hardware, large-scale infrastructure, DevOps, and configuration-management expertise.
Top Skills:
AnsibleArmDevOpsDistributed SystemsKvm/QemuLinuxLinux KernelNested VirtualizationNvidia GraceObservability InfrastructureSaltstack
Cloud • Security • Software • Cybersecurity
Architect, build, and support reliable network infrastructure and automation for Akamai’s distributed cloud platform. Develop Bash and Python tooling, establish deployment standards, define SLOs, mentor engineers, and troubleshoot complex network issues. The role requires expertise in large-scale distributed systems, TCP/IP, BGP, Linux networking, configuration management, CI/CD, and open-source networking software, with participation in on-call rotations.
Top Skills:
AnsibleArgocdBashBgpBirdChefFirewallsFrrGithub ActionsGoGobgpJenkinsLinux NetworkingLoad BalancingPuppetPythonRustSaltstackSlack BotsTcp/Ip
Software
Own reliability, performance, scalability, and operational standards across on-premises, private-cloud, and AWS environments. Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling; define SLOs, SLIs, error budgets, and readiness standards. Lead incident response, root-cause analysis, disaster-recovery readiness, release coordination, and preventive automation. Mentor engineers, coach teams on operational practices, coordinate on-call coverage, and ensure infrastructure meets security and compliance requirements.
Top Skills:
AnsibleAuto ScalingAWSBashCi/CdClaudeCloudwatchCrowdstrikeDatadogDistributed SystemsDnsDockerEc2EcsGitGithub CopilotGitlabGitopsIamKubernetesLinuxLoad BalancingOracle LinuxPythonQualysRapid7RhelS3Tcp/IpTerraformVpcWireshark
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills:
AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Artificial Intelligence • Hardware • Software • Semiconductor
Develop kernel-centric reliability solutions for AI compute clusters and production services. Responsibilities include debugging failures, building diagnostic tools, supporting incident response, performing root-cause analysis, improving kernel and software reliability, and collaborating with systems, hardware, ASIC, and architecture teams on reliability-focused designs.
Top Skills:
CC++Core Dump HandlingDebuggersDistributed ProgrammingGpusParallel ProgrammingProfilersPythonSanitizersTracing
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve large-scale Kubernetes and GPU clusters across public and private clouds. Build automation, observability, capacity-management, and reliability systems; define SLOs and SLIs; support production launches; lead incident triage and root-cause analysis; conduct blameless postmortems; and participate in on-call support for AI workloads.
Top Skills:
AirflowAnsibleArgo WorkflowsAWSAws Step FunctionsCadenceChefCudaElk StackGoGoogle Cloud PlatformGrafanaKubernetesKubevirtLightstepLinuxAzureNcclNvidia Dgx CloudNvidia DynamoOpentelemetryOracle Cloud InfrastructurePrometheusPuppetPythonPyTorchSglangSplunkTcp/IpTemporalTensorrt-LlmTerraformVllm
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results




























.jpg)


