Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in San Francisco, CA
Artificial Intelligence • Logistics • Robotics • Software
Own reliability across cloud, edge, and on-site deployments. Build observability, monitoring, and alerting. Define incident response and on-call processes, improve deployment workflows, diagnose infra/network/distributed-system issues, and make deployments repeatable and scalable.
Top Skills:
AWSAzureGCPGrafanaKafkaKubernetesLinuxOpentelemetryPrometheusRtspSecure TunnelsVpnWebrtc
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills:
AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
Big Data • Cloud • Marketing Tech • Social Impact • Software
The Senior Staff Site Reliability Engineer at LiveRamp will define the SRE strategy, oversee critical automation, and lead operational excellence in a global infrastructure, influencing architectural decisions and mentoring teams.
Top Skills:
Aws)CassandraCircleCICloud Security (GcpDynamoDBGoJenkinsKubernetesPythonScylladbSinglestoreTerraform
Cloud
The Site Reliability Engineer will manage Kubernetes platforms, optimize AWS cloud infrastructure, ensure high availability, and automate deployment while handling troubleshooting and security compliance.
Top Skills:
AWSBashCi/CdCloudwatchElk StackGoGrafanaHelmIstioKubernetesPrometheusPythonTerraform
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills:
AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Cloud
The role involves building and managing observability infrastructure in GCP, automating deployments, and optimizing data processes for high reliability.
Top Skills:
GkeGoGCPGrafanaKubernetesOpentelemetryPythonRubySplunkTerraform
Security • Software
Build, scale, and maintain Tenable's cloud-based vulnerability management platform for private and U.S. government customers. Troubleshoot escalations, automate deployments and monitoring, collaborate across engineering teams, document operational procedures, participate in on-call rotation, remediate infrastructure/container security issues, and contribute to reliability, standardization, and production deployments.
Top Skills:
BashContainersDockerKubernetesMicroservicesNode.jsPythonTerraform
Reposted 13 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Payments • Software • Financial Services
Lead SRE responsible for reliability strategy, architecting resilient AWS/Kubernetes infrastructure, building observability, driving incident response and postmortems, improving deployment safety and automation, mentoring engineers, and partnering across product and engineering to scale platform reliability.
Top Skills:
AWSCi/CdDockerEvent-Driven ArchitecturesFastapiInfrastructure As CodeKubernetes (Eks)LogsObservability (MetricsPostgresql (Rds)PythonTraces)TypescriptVue
Artificial Intelligence • Information Technology • Software
Lead end-to-end platform reliability: define SLIs/SLOs, harden production architecture, ensure Kubernetes runtime and queue safety, run incident command for Sev1/Sev2, own observability/on-call/runbooks, and gate risky releases while delivering a prioritized reliability roadmap.
Top Skills:
BullmqKoaKubernetesNode.jsPostgraphilePostgresReactRedisTypescript
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own reliability for a live fleet of Linux-based edge sensors and cloud infrastructure. Triage and recover field hardware, perform SSH-based diagnostics, build fleet management and OTA systems, implement observability and alerting, automate operational tasks, develop runbooks, and participate in on-call rotations to prevent and resolve incidents.
Top Skills:
AWSBashCDnsDockerFirewallsGoIamKubernetesLinuxPythonRustSshVpn
Information Technology • Mobile • Software
As a Site Reliability Engineer, you'll ensure system reliability and scalability, automate processes, optimize performance, and collaborate on system design.
Top Skills:
AWSAzureBashCloudFormationDatadogDockerElkGoGoogle Cloud PlatformGrafanaHelmKubernetesNew RelicPrometheusPulumiPythonTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Software
Own the reliability and performance of backend systems at Gamma, building automation and tooling while leading incident response and improving system stability.
Top Skills:
AWSCloudFormationDockerGoKafkaKubernetesNode.jsPythonTerraformTypescript
Artificial Intelligence • Information Technology • Software • Automation
Own US PST coverage for releases and incidents as the first SRE; bridge infrastructure and code by working with Kubernetes, Terraform, and AWS and patching Elixir when needed; lead incident response and post-mortems; define SLOs and observability; author runbooks and support HIPAA-aligned compliance for a regulated medical-device platform.
Top Skills:
AWSElixirKubernetesTerraform
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills:
AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Cloud • Software • Analytics
Join Arista Networks as a Site Reliability Engineer to manage CloudVision service reliability, scalability, and stability in a FedRAMP environment, focusing on areas like architecture, security, and performance optimization.
Top Skills:
AnsibleBashGCPGkeGoKubernetesPulumiPython
Artificial Intelligence • Software • Generative AI • Automation
Operate and harden Blitzy's self-hosted, Kubernetes-based AI platform inside customer-controlled secure cloud environments. Own deployments, upgrades, capacity planning, observability, incident response, and customer-facing technical coordination while championing security and feeding operational learnings back into the product roadmap.
Top Skills:
Alerting)BashCloud (Aws/Gcp/Azure)Container OrchestrationGoInfrastructure-As-CodeKubernetesMetricsObservability (LoggingPulumiPythonTerraformTracing
Other
Lead architecture and delivery of highly available, resilient cloud and on‑prem systems for the Password Safe platform. Own platform engineering, CI/CD pipelines, IaC/GitOps, release orchestration, observability (metrics/logs/traces), chaos engineering, SLO/SLI definition, and core services. Mentor engineers, define SRE strategy, and drive reliability, security, and automation improvements.
Top Skills:
AnsibleApi GatewaysAWSAzureBlue Green DeploymentsC#CachesCanary DeploymentsChaos EngineeringCi/CdConfiguration ManagementDatadogDevsecopsDockerGitopsGoGrafana CloudJavaKubernetesLinuxOpentelemetryOpentofuSecrets ManagementService MeshTerraformWindows
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and deploy AI/ML tools and LLM/agent-based systems for GeForce NOW SRE. Transform production signals, metrics, and logs into actionable intelligence, automate root-cause analysis, manage large-scale data pipelines, and operate ML tooling on Kubernetes and AWS with monitoring and visualization (Grafana).
Top Skills:
Agent-Based SystemsAWSGoGrafanaKubernetesLlmsPython
Information Technology • Legal Tech
The Senior Technology Site Reliability Engineer is responsible for maintaining and optimizing infrastructure and applications, ensuring reliability and performance while automating processes and collaborating with teams.
Top Skills:
AWSChefDatadogGoGrafanaJavaPrometheusPuppetPythonSaltTerraform
Artificial Intelligence • Logistics • Software
The Site Reliability Engineer will enhance operational resilience, ensuring system stability, observability, and debugging workflows for complex failures while improving developer focus and uptime.
Top Skills:
DatadogGoPrometheusPythonSentry
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Staff Software Engineer in Site Reliability, you'll manage infrastructure for reliability and scalability, lead incident management, and automate operational tasks.
Top Skills:
AWSAzureBashCloudFormationDatadogGCPGoIncidentioPagerdutyPulumiPythonSentryTerraform
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Software Engineer in Site Reliability, you will ensure the reliability and performance of our AI platform through automation and strategic infrastructure management.
Top Skills:
AWSAzureBashCloudFormationDatadogGCPGoKubernetesPagerdutyPythonSentryTerraform
Artificial Intelligence • Software
As a Software Engineer on the Site Reliability team, you'll ensure system reliability, scalability, and observability while partnering with engineering teams and improving incident management processes.
Top Skills:
AWSCi/Cd ToolingContainer OrchestrationDatadogGrafanaPrometheusTerraform
Artificial Intelligence • Machine Learning • Robotics • Software • Transportation • Design • Manufacturing
The Staff Site Reliability Engineer will lead source control strategy, manage Git-based monorepo operations, improve developer productivity, and oversee migrations to GitHub Cloud.
Top Skills:
BazelBuckBuildkiteGerritGithub ActionsGithub CloudGithub EnterpriseGitlab CiJenkinsPulumiReviewableTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results















_1.png)















