Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in San Francisco, CA
Information Technology • Software • Big Data Analytics
The Site Reliability Engineer will design, analyze, and troubleshoot large-scale distributed systems, focusing on operating systems and performance tuning.
Top Skills:
ApacheJava
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills:
ArgocdAWSDatadogGitGoKubernetesPythonTerraform
Financial Services
Provide frontline desktop support for employees (remote and in-person), triage and resolve hardware, Windows, application, phone, and market-data feed issues, manage tickets, perform firmware/patch deployments, and collaborate with IT teams. Support trading desk/C-suite users and maintain endpoint security and configuration management.
Top Skills:
Active DirectoryBiometric DevicesBloombergCisco Phone SystemsCisco PhonesData EncryptionEndpoint ManagerFidessaFirmware UpdatesGlobal RelayIceMicrosoft Office/Office 365Ms-900OnedrivePatch ManagementPrintersRedi+ScannersServicenowSoftphonesSpyware/Malware ToolsSystem Center Configuration ManagerThomson ReutersTrading TurretsVpnWifiWindows 10Windows 11Zoom
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills:
AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Artificial Intelligence
The Deployment Engineer will build and operate AI inference clusters, ensure scalable deployments, optimize allocation, and maintain infrastructure. Responsibilities include software updates, telemetry development, and collaborative improvements with teams.
Top Skills:
DockerGrafanaInfluxdbK8SLinuxPrometheusPython
Cloud • Security • Software • Cybersecurity
Build and operate the Veeam Data Cloud GOV environment: map systems, write runbooks, define SLIs/SLOs, run incident response, close observability gaps, automate deployments and support fleet management while working across security and compliance constraints.
Top Skills:
Api ManagementApplication InsightsArgocdAws CloudformationAzureAzure Arm TemplatesAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
Artificial Intelligence • Software • Generative AI
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating cloud-native systems, owning production reliability and incident response, building observability and automation, defining SLOs/SLIs, driving postmortems, and partnering with product and engineering teams to improve operational maturity.
Top Skills:
AWSAzureGCPGoJavaKubernetesPython
10 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Artificial Intelligence • Fintech • Machine Learning • Software • Financial Services
Lead product lifecycle for internal PaaS and Data-as-a-Service offerings focused on SRE and data engineering. Drive system reliability (99.99% availability), scaling strategies, GCP cost optimization, BigQuery/table and ETL performance, AI/ML enablement, and data governance/compliance across cross-functional teams.
Top Skills:
AIAWSBigQueryBusiness IntelligenceCloud IamData LakeData WarehouseEltETLFinopsGCPGkeGoogle Cloud PlatformKubernetesMachine LearningSite Reliability EngineeringSreTerraform
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills:
Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills:
ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills:
AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Information Technology • Marketing Tech • Social Media
Own and maintain reliability, performance, scalability, and capacity of massive Ceph storage clusters. Troubleshoot distributed storage issues, build automation (Python/Shell, SaltStack, Ansible), define observability (SLIs/SLOs, PromQL/LogQL, Grafana), and lead lifecycle initiatives like upgrades, expansions, and migrations.
Top Skills:
AnsibleCeph (RadosCephfs)ChefGrafanaLinuxLogqlPromqlPuppetPythonRbdRgwSaltstackShell
Security • Software • Cybersecurity • Automation
As a Senior Site Reliability Engineer, you will enhance the reliability of Drata’s product teams through automation, architecture reviews, and operational excellence using cloud-native technologies.
Top Skills:
AiopsAWSBashDatadogDockerGitGithub ActionsKubernetesLinuxMySQLPythonTerraform
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills:
DhcpDnsLinuxNtpPython
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills:
Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
Software
Lead improvements in reliability, performance, scalability, capacity, and observability through automation and tooling. Partner with developers to design infrastructure and monitoring, define SLIs/SLOs, conduct load and performance testing, participate in incident response and root cause analysis, and manage monitoring services for production systems.
Top Skills:
DockerGoGradleJavaKubernetesOpentelemetrySpring BootTerraform
Big Data • Energy • Big Data Analytics
The Staff Site Reliability Engineer will lead in designing and maintaining cloud infrastructure on GCP, drive IaC strategy, manage Kubernetes operations, ensure security compliance, and mentor engineers.
Top Skills:
BashGoGoogle Cloud PlatformGrafanaKubernetesOpentelemetryPostgresPythonTerraform
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills:
AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills:
ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Information Technology • Internet of Things • Software • Virtual Reality
Lead reliability, availability, and resiliency strategies for large-scale systems, drive operational excellence, and provide technical mentorship across engineering teams.
Top Skills:
AWSCi/CdJavaMongoDBRabbitMQZookeeper
Healthtech
Build, operate, and scale AWS cloud infrastructure and Kubernetes workloads using Terraform and Helm. Improve observability, define SLIs/SLOs, automate deployments and incident response, support on-call rotation, and implement security and compliance (HIPAA, SOC 2) best practices while partnering with product and engineering teams.
Top Skills:
AWSCi/CdEvent SourcingHelmKubernetesLinuxMonitoring/Logging/TracingNetworkingTerraform
Fintech • Payments
The Senior Staff SRE leads reliability engineering initiatives, drives operational excellence, mentors staff, and influences architecture to enhance system reliability and performance.
Top Skills:
Ai/MlAWSAzureDockerElk StackGCPGrafanaKubernetesMySQLNoSQLPostgresSplunk
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills:
Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Hardware • Machine Learning • Security • Software
Design and maintain developer experience tooling and CI/CD pipelines, manage self-service deployment tools (secrets, rollbacks), partner with cloud teams to troubleshoot production systems, participate in on-call rotation, and write production-grade automation and tests in Go and TypeScript to improve developer velocity and reliability.
Top Skills:
AWSGithub ActionsGoHelmKubernetesTerraformTypescript
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results









.png)







%20(1).png)


.png)










