Top Senior Site Reliability Engineer Jobs in San Francisco, CA

Reposted 11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
Mid level
Mid level
Information Technology • Software • Big Data Analytics
The Site Reliability Engineer will design, analyze, and troubleshoot large-scale distributed systems, focusing on operating systems and performance tuning.
Top Skills: ApacheJava
Reposted 11 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills: ArgocdAWSDatadogGitGoKubernetesPythonTerraform
Reposted 11 Days AgoSaved
In-Office
San Francisco Bay Area, CA
85K-115K Annually
Mid level
85K-115K Annually
Mid level
Financial Services
Provide frontline desktop support for employees (remote and in-person), triage and resolve hardware, Windows, application, phone, and market-data feed issues, manage tickets, perform firmware/patch deployments, and collaborate with IT teams. Support trading desk/C-suite users and maintain endpoint security and configuration management.
Top Skills: Active DirectoryBiometric DevicesBloombergCisco Phone SystemsCisco PhonesData EncryptionEndpoint ManagerFidessaFirmware UpdatesGlobal RelayIceMicrosoft Office/Office 365Ms-900OnedrivePatch ManagementPrintersRedi+ScannersServicenowSoftphonesSpyware/Malware ToolsSystem Center Configuration ManagerThomson ReutersTrading TurretsVpnWifiWindows 10Windows 11Zoom
Reposted 12 Days AgoSaved
In-Office
San Francisco Bay Area, CA
116K-174K Annually
Senior level
116K-174K Annually
Senior level
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills: AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Reposted 3 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence
The Deployment Engineer will build and operate AI inference clusters, ensure scalable deployments, optimize allocation, and maintain infrastructure. Responsibilities include software updates, telemetry development, and collaborative improvements with teams.
Top Skills: DockerGrafanaInfluxdbK8SLinuxPrometheusPython
4 Days AgoSaved
Remote
San Francisco Bay Area, CA
152K-253K Annually
Senior level
152K-253K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Build and operate the Veeam Data Cloud GOV environment: map systems, write runbooks, define SLIs/SLOs, run incident response, close observability gaps, automate deployments and support fleet management while working across security and compliance constraints.
Top Skills: Api ManagementApplication InsightsArgocdAws CloudformationAzureAzure Arm TemplatesAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
Reposted 13 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Software • Generative AI
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating cloud-native systems, owning production reliability and incident response, building observability and automation, defining SLOs/SLIs, driving postmortems, and partnering with product and engineering teams to improve operational maturity.
Top Skills: AWSAzureGCPGoJavaKubernetesPython
10 Days AgoSaved
Remote
San Francisco Bay Area, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 4 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Artificial Intelligence • Fintech • Machine Learning • Software • Financial Services
Lead product lifecycle for internal PaaS and Data-as-a-Service offerings focused on SRE and data engineering. Drive system reliability (99.99% availability), scaling strategies, GCP cost optimization, BigQuery/table and ETL performance, AI/ML enablement, and data governance/compliance across cross-functional teams.
Top Skills: AIAWSBigQueryBusiness IntelligenceCloud IamData LakeData WarehouseEltETLFinopsGCPGkeGoogle Cloud PlatformKubernetesMachine LearningSite Reliability EngineeringSreTerraform
Reposted 4 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills: Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
Reposted 5 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills: ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Reposted 5 Days AgoSaved
Remote
San Francisco Bay Area, CA
154K-231K Annually
Senior level
154K-231K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills: AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 5 Days AgoSaved
Remote
San Francisco Bay Area, CA
128K-192K Annually
Senior level
128K-192K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Own and maintain reliability, performance, scalability, and capacity of massive Ceph storage clusters. Troubleshoot distributed storage issues, build automation (Python/Shell, SaltStack, Ansible), define observability (SLIs/SLOs, PromQL/LogQL, Grafana), and lead lifecycle initiatives like upgrades, expansions, and migrations.
Top Skills: AnsibleCeph (RadosCephfs)ChefGrafanaLinuxLogqlPromqlPuppetPythonRbdRgwSaltstackShell
Reposted 20 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
167K-226K Annually
Senior level
167K-226K Annually
Senior level
Security • Software • Cybersecurity • Automation
As a Senior Site Reliability Engineer, you will enhance the reliability of Drata’s product teams through automation, architecture reviews, and operational excellence using cloud-native technologies.
Top Skills: AiopsAWSBashDatadogDockerGitGithub ActionsKubernetesLinuxMySQLPythonTerraform
Reposted 6 Days AgoSaved
Remote
San Francisco Bay Area, CA
100K-140K Annually
Mid level
100K-140K Annually
Mid level
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills: DhcpDnsLinuxNtpPython
7 Days AgoSaved
Remote
San Francisco Bay Area, CA
75K-90K Annually
Mid level
75K-90K Annually
Mid level
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills: Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
7 Days AgoSaved
Remote
San Francisco Bay Area, CA
175K-185K Annually
Senior level
175K-185K Annually
Senior level
Software
Lead improvements in reliability, performance, scalability, capacity, and observability through automation and tooling. Partner with developers to design infrastructure and monitoring, define SLIs/SLOs, conduct load and performance testing, participate in incident response and root cause analysis, and manage monitoring services for production systems.
Top Skills: DockerGoGradleJavaKubernetesOpentelemetrySpring BootTerraform
Reposted 16 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
136K-180K Annually
Senior level
136K-180K Annually
Senior level
Big Data • Energy • Big Data Analytics
The Staff Site Reliability Engineer will lead in designing and maintaining cloud infrastructure on GCP, drive IaC strategy, manage Kubernetes operations, ensure security compliance, and mentor engineers.
Top Skills: BashGoGoogle Cloud PlatformGrafanaKubernetesOpentelemetryPostgresPythonTerraform
17 Days AgoSaved
In-Office
San Francisco Bay Area, CA
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills: AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
17 Days AgoSaved
In-Office
San Francisco Bay Area, CA
157K-239K Annually
Senior level
157K-239K Annually
Senior level
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills: ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Reposted 7 Days AgoSaved
Remote
San Francisco Bay Area, CA
131K-185K Annually
Senior level
131K-185K Annually
Senior level
Information Technology • Internet of Things • Software • Virtual Reality
Lead reliability, availability, and resiliency strategies for large-scale systems, drive operational excellence, and provide technical mentorship across engineering teams.
Top Skills: AWSCi/CdJavaMongoDBRabbitMQZookeeper
8 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
140K-150K Annually
Mid level
140K-150K Annually
Mid level
Healthtech
Build, operate, and scale AWS cloud infrastructure and Kubernetes workloads using Terraform and Helm. Improve observability, define SLIs/SLOs, automate deployments and incident response, support on-call rotation, and implement security and compliance (HIPAA, SOC 2) best practices while partnering with product and engineering teams.
Top Skills: AWSCi/CdEvent SourcingHelmKubernetesLinuxMonitoring/Logging/TracingNetworkingTerraform
Reposted 17 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
160K-179K Annually
Senior level
160K-179K Annually
Senior level
Fintech • Payments
The Senior Staff SRE leads reliability engineering initiatives, drives operational excellence, mentors staff, and influences architecture to enhance system reliability and performance.
Top Skills: Ai/MlAWSAzureDockerElk StackGCPGrafanaKubernetesMySQLNoSQLPostgresSplunk
Reposted 18 Days AgoSaved
In-Office
San Francisco Bay Area, CA
173K-224K Annually
Senior level
173K-224K Annually
Senior level
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills: Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
9 Days AgoSaved
Remote
San Francisco Bay Area, CA
140K-165K Annually
Senior level
140K-165K Annually
Senior level
Hardware • Machine Learning • Security • Software
Design and maintain developer experience tooling and CI/CD pipelines, manage self-service deployment tools (secrets, rollbacks), partner with cloud teams to troubleshoot production systems, participate in on-call rotation, and write production-grade automation and tests in Go and TypeScript to improve developer velocity and reliability.
Top Skills: AWSGithub ActionsGoHelmKubernetesTerraformTypescript
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account