Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in San Francisco Bay Area, CA
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead the establishment and maturation of the SRE function across cloud infrastructure and platform services. Define reliability targets, build observability and automation, lead incident response and root-cause analysis, improve resilience and recovery, and reduce operational toil. Partner with product and platform teams on reliability-focused design, establish incident practices, manage the SRE roadmap, and mentor engineers across Cloud Engineering and Reliability.
Top Skills:
AWSCloud-Native ObservabilityGoInfrastructure As CodeKubernetesPython
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills:
Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills:
AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 16 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 18 Days AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 20 Days AgoSaved
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills:
ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills:
Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
Reposted 13 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 14 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Lead design and delivery of scalable cloud infrastructure for the Spend product. Embed with development teams to drive reliability, performance, observability, incident response, and automation. Own SLOs, runbooks, DevOps metrics, and collaborate with central DevOps and security teams to ensure compliance and resilience. Lead infrastructure projects including new service launches, data centre migrations, and modernising data pipelines.
Top Skills:
Analytics PipelinesAWSData StreamingDevOpsGCPIncident ResponseKubernetesObservabilitySlosSre
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor cloud security posture with Wiz, and lead major-incident response as Incident Commander. Drive remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and develop automated runbooks for compliant, secure platforms.
Top Skills:
AWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWizZero-Trust Security
Fintech • Professional Services • Software
Build and improve scalable, resilient backend systems across product teams. Responsibilities include implementing observability with OpenTelemetry, establishing SLOs, owning incident response processes, mentoring teams on secure and reliable coding practices, and applying automated testing. The role collaborates with engineering leadership and cross-functional partners on shared infrastructure, performance, and distributed-systems challenges. This is a hybrid position requiring at least three office days per week in Los Angeles or San Francisco.
Top Skills:
AWSClaudeDebeziumDistributed SystemsDockerGitlab Ci/CdHelmJavaKafkaKubernetesMicroservicesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Fintech • Professional Services • Software
Own performance, resilience, observability, and incident response for scalable backend systems across product teams. Mentor engineers on secure, reliable coding practices; implement OpenTelemetry, logging, SLOs, and automated testing; and contribute shared infrastructure solutions. The role requires expertise in Java, Spring Boot, REST APIs, microservices, distributed systems, SQL, Docker, Kubernetes, and Helm, with AWS, CI/CD, infrastructure tooling, and financial services experience preferred.
Top Skills:
AWSClaudeDebeziumDockerGitlab Ci/CdHelmJavaKafkaKubernetesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills:
AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills:
Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills:
Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 22 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Fintech • Payments • Financial Services
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
Top Skills:
Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayDastDatadogDockerGitlabGrafanaIamJavaKubernetesLlmsNew RelicNode.jsOwasp Top 10PythonSastSplunkTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Designs, deploys, and operates reliable, scalable cloud and network infrastructure; automates platforms; monitors SLOs/SLAs; leads on-call incident response, postmortems, DR planning, and security controls to improve service reliability.
Top Skills:
AutomationCloud InfrastructureDisaster RecoveryFedramp HighIl-5Monitoring And AlertingNetworkingOn-Call Incident ManagementSecurity ControlsSre PracticesStorage Systems
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Bay Area, CA Companies Hiring SRE Engineers
See AllPopular San Francisco Bay Area, CA Engineering Job Searches
Engineering Jobs in San Francisco Bay Area, CA
Software Engineer Jobs in San Francisco Bay Area, CA
Android Developer Jobs in San Francisco Bay Area, CA
C# Jobs in San Francisco Bay Area, CA
C++ Jobs in San Francisco Bay Area, CA
DevOps Jobs in San Francisco Bay Area, CA
Front End Developer Jobs in San Francisco Bay Area, CA
Golang Jobs in San Francisco Bay Area, CA
Hardware Engineer Jobs in San Francisco Bay Area, CA
iOS Developer Jobs in San Francisco Bay Area, CA
Java Developer Jobs in San Francisco Bay Area, CA
Javascript Jobs in San Francisco Bay Area, CA
Linux Jobs in San Francisco Bay Area, CA
Engineering Manager Jobs in San Francisco Bay Area, CA
.NET Developer Jobs in San Francisco Bay Area, CA
PHP Developer Jobs in San Francisco Bay Area, CA
Python Jobs in San Francisco Bay Area, CA
QA Jobs in San Francisco Bay Area, CA
Ruby Jobs in San Francisco Bay Area, CA
Salesforce Developer Jobs in San Francisco Bay Area, CA
Scala Jobs in San Francisco Bay Area, CA
Application Engineer Jobs in San Francisco, CA
Associate Software Engineer Jobs in San Francisco Bay Area, CA
Automation Engineer Jobs in San Francisco Bay Area, CA
AWS Engineer Jobs in San Francisco, CA
Backend Engineer Jobs in San Francisco, CA
Cloud Engineer Jobs in San Francisco Bay Area, CA
Controls Engineer Jobs in San Francisco Bay Area, CA
CTO Jobs in San Francisco Bay Area, CA
Design Engineer Jobs in San Francisco Bay Area, CA
DevOps Engineer Jobs in San Francisco Bay Area, CA
Director of Engineering Jobs in San Francisco, CA
Electrical Engineering Jobs in San Francisco Bay Area, CA
Embedded Software Engineer Jobs in San Francisco Bay Area, CA
Field Engineer Jobs in San Francisco Bay Area, CA
Firmware Engineer Jobs in San Francisco, CA
Full-Stack Engineer Jobs in San Francisco, CA
Game Engineer Jobs in San Francisco, CA
Infrastructure Engineer Jobs in San Francisco, CA
Manufacturing Engineer Jobs in San Francisco Bay Area, CA
Mechanical Design Engineer Jobs in San Francisco Bay Area, CA
Mechanical Engineering Jobs in San Francisco Bay Area, CA
Mechatronics Engineering Jobs in San Francisco Bay Area, CA
Network Engineer Jobs in San Francisco Bay Area, CA
Platform Engineer Jobs in San Francisco, CA
Process Engineer Jobs in San Francisco Bay Area, CA
Project Engineer Jobs in San Francisco Bay Area, CA
QA Engineer Jobs in San Francisco Bay Area, CA
Robotics Engineer Jobs in San Francisco Bay Area, CA
Security Engineer Jobs in San Francisco Bay Area, CA
Software Engineering Manager Jobs in San Francisco, CA
Software Test Engineer Jobs in San Francisco, CA
SRE Engineer Jobs in San Francisco Bay Area, CA
Systems Engineer Jobs in San Francisco Bay Area, CA
VP of Engineering Jobs in San Francisco, CA
All Filters
Total selected ()
No Results
No Results


.png)













.png)











