Top SRE Engineer Jobs in San Francisco Bay Area, CA

2 Days AgoSaved
In-Office
San Francisco Bay Area, CA
220K-340K Annually
Senior level
220K-340K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead the establishment and maturation of the SRE function across cloud infrastructure and platform services. Define reliability targets, build observability and automation, lead incident response and root-cause analysis, improve resilience and recovery, and reduce operational toil. Partner with product and platform teams on reliability-focused design, establish incident practices, manage the SRE roadmap, and mentor engineers across Cloud Engineering and Reliability.
Top Skills: AWSCloud-Native ObservabilityGoInfrastructure As CodeKubernetesPython
8 Days AgoSaved
Easy Apply
Hybrid
San Francisco Bay Area, CA
Easy Apply
150K-220K Annually
Expert/Leader
150K-220K Annually
Expert/Leader
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills: Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
3 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
191K-226K Annually
Senior level
191K-226K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills: AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 14 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
214K-260K Annually
Senior level
214K-260K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills: AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 16 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 18 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills: AWSGoKubernetesPuppetPythonTerraform
Reposted 18 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 20 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
22 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills: Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
Reposted 13 Days AgoSaved
Remote
San Francisco Bay Area, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 14 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Reposted 18 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 28 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
160K-250K Annually
Senior level
160K-250K Annually
Senior level
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Lead design and delivery of scalable cloud infrastructure for the Spend product. Embed with development teams to drive reliability, performance, observability, incident response, and automation. Own SLOs, runbooks, DevOps metrics, and collaborate with central DevOps and security teams to ensure compliance and resilience. Lead infrastructure projects including new service launches, data centre migrations, and modernising data pipelines.
Top Skills: Analytics PipelinesAWSData StreamingDevOpsGCPIncident ResponseKubernetesObservabilitySlosSre
Reposted 3 Hours AgoSaved
In-Office
San Francisco Bay Area, CA
Senior level
Senior level
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor cloud security posture with Wiz, and lead major-incident response as Incident Commander. Drive remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and develop automated runbooks for compliant, secure platforms.
Top Skills: AWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWizZero-Trust Security
4 Hours AgoSaved
In-Office
San Francisco Bay Area, CA
200K-240K Annually
Senior level
200K-240K Annually
Senior level
Fintech • Professional Services • Software
Build and improve scalable, resilient backend systems across product teams. Responsibilities include implementing observability with OpenTelemetry, establishing SLOs, owning incident response processes, mentoring teams on secure and reliable coding practices, and applying automated testing. The role collaborates with engineering leadership and cross-functional partners on shared infrastructure, performance, and distributed-systems challenges. This is a hybrid position requiring at least three office days per week in Los Angeles or San Francisco.
Top Skills: AWSClaudeDebeziumDistributed SystemsDockerGitlab Ci/CdHelmJavaKafkaKubernetesMicroservicesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
4 Hours AgoSaved
In-Office
San Francisco Bay Area, CA
203K-275K Annually
Senior level
203K-275K Annually
Senior level
Fintech • Professional Services • Software
Own performance, resilience, observability, and incident response for scalable backend systems across product teams. Mentor engineers on secure, reliable coding practices; implement OpenTelemetry, logging, SLOs, and automated testing; and contribute shared infrastructure solutions. The role requires expertise in Java, Spring Boot, REST APIs, microservices, distributed systems, SQL, Docker, Kubernetes, and Helm, with AWS, CI/CD, infrastructure tooling, and financial services experience preferred.
Top Skills: AWSClaudeDebeziumDockerGitlab Ci/CdHelmJavaKafkaKubernetesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
YesterdaySaved
In-Office or Remote
San Francisco Bay Area, CA
320K-485K Annually
Senior level
320K-485K Annually
Senior level
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills: AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
Reposted 20 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills: Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
195K-258K Annually
Senior level
195K-258K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills: Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
21 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 22 Days AgoSaved
Easy Apply
Remote or Hybrid
San Francisco Bay Area, CA
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted 24 Days AgoSaved
Remote or Hybrid
San Francisco Bay Area, CA
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
7 Days AgoSaved
In-Office
San Francisco Bay Area, CA
147K-234K Annually
Senior level
147K-234K Annually
Senior level
Fintech • Payments • Financial Services
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
Top Skills: Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayDastDatadogDockerGitlabGrafanaIamJavaKubernetesLlmsNew RelicNode.jsOwasp Top 10PythonSastSplunkTerraform
Reposted 8 Days AgoSaved
In-Office
San Francisco Bay Area, CA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Designs, deploys, and operates reliable, scalable cloud and network infrastructure; automates platforms; monitors SLOs/SLAs; leads on-call incident response, postmortems, DR planning, and security controls to improve service reliability.
Top Skills: AutomationCloud InfrastructureDisaster RecoveryFedramp HighIl-5Monitoring And AlertingNetworkingOn-Call Incident ManagementSecurity ControlsSre PracticesStorage Systems
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account