Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in San Francisco Bay Area, CA
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills:
Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Information Technology • Consulting
As a Senior Staff Site Reliability Engineer, you will lead the SRE team, advocate best practices, ensure resilience in cloud architecture, and mentor team members.
Top Skills:
ArgocdCircleCIGoogle Cloud PlatformKubernetesPulumiTerraformTypescript
Enterprise Web • Information Technology • Software
As a Platform Engineer, you will enhance reliability and performance, design operational processes, and build monitoring systems while collaborating with a talented team.
Top Skills:
AIAssistantsBackendDeveloper ToolsFrontendInfrastructureMcpsMonitoring SystemsSkills
Artificial Intelligence • Software
As a Site Reliability Engineer at Mercor, you will ensure production reliability, develop SRE function, and collaborate with engineering teams to maintain system performance.
Top Skills:
AWSKubernetesSpaceliftTerraform
Software
Join a passionate team to enhance reliability and performance of the AI control plane, manage deployments, and respond to production incidents while ensuring service quality for customers.
Top Skills:
Ai Control PlaneDeveloper ToolsInfrastructure
Aerospace • Artificial Intelligence
The Site Reliability Engineer will architect and manage ground infrastructure for satellite systems, ensuring high availability, automating deployments, and optimizing data management systems.
Top Skills:
AnsibleAWSAzureC++CloudFormationEksElkGCPGrafanaHelmKubernetesPrometheusPythonTerraform
Fintech • Payments • Software • Financial Services
Senior SRE responsible for ensuring platform scalability, reliability, and runtime efficiency on AWS. Own CI/CD and GitHub repo workflows, lead incident response and post-mortems, implement observability/monitoring and logging, and collaborate cross-border using bilingual Mandarin and English.
Top Skills:
AlertingAWSCi/CdDeployment PipelinesGitGithub ActionsLoggingMonitoringObservabilityScripting
Reposted 3 Days AgoSaved
Artificial Intelligence • Information Technology • Software • Automation
Lead technical vision as a principal engineer, either managing teams or driving cross-team initiatives. Design and architect cloud infrastructure, networking, and security; define authentication/authorization patterns; architect and operate Kubernetes deployments; and implement infrastructure-as-code using tools like Terraform, CloudFormation, Ansible, or Puppet.
Top Skills:
AnsibleAWSCloudFormationGCPIamKubernetesPuppetRbacSecurity GroupsTerraform
Artificial Intelligence • Healthtech • Software • Automation
Design, build, and operate Optura's multi-cloud, HIPAA-aware platform: run Kubernetes across cloud and customer on-prem/air-gapped environments, create unified deployment tooling (Helm/operators/GitOps), own SLOs/capacity/incident response, drive reliability, implement identity/networking/security controls, and build IaC/GitOps patterns in partnership with product and security teams.
Top Skills:
AksArgo CdAWSAzureBackstageCluster ApiCrossplaneDistributed TracingEksGCPGitopsGkeGoGrafanaHelmKmsKubernetesMtlsOidcOpenshiftOpentelemetryOperatorsPrometheusPulumiPythonRancherReplicatedSecrets ManagementService MeshTalosTerraformVpc
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for Autodesk GovCloud services by deploying, operating, and automating production systems. Define SLOs/SLIs, build observability and automation, run incident response and on-call rotation, ensure compliance (FedRAMP), perform resilience testing and toil reduction, and collaborate across engineering, security, and platform teams to improve service reliability and operability.
Top Skills:
APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDnsDynatraceFedrampGoIl4Il5Infrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Software
As an AI Support Engineer, you'll manage support requests, resolve user issues, optimize ML models, and contribute to product development.
Top Skills:
Tensorrt
Cloud
Design, build, and maintain cloud platform services for sensitive federal missions. Operate and monitor air-gapped, secure environments; build CI/CD pipelines without internet; maintain SLOs/SLIs; own runbooks and incident response; support POA&M and Authority to Operate activities; and deliver internal platform enablement while ensuring strict compliance and security.
Top Skills:
Aws VpcsBgpCi/CdCloudwatchContainersEcs FargateEksGrafanaIpsecLinuxPythonSplunkTerraformTgwsVpc Endpoints
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Information Technology
The Site Reliability Engineer will drive reliability for the Tinker platform, focusing on incident response, monitoring, and ensuring system resilience while collaborating across teams.
Top Skills:
Cloud InfrastructureKubernetes
Artificial Intelligence • Information Technology • Software
Build and operate the core infrastructure for Arena's online evaluation systems: design low-latency, high-reliability APIs and gateways, implement enterprise-grade features (rate limiting, auth, metering, audit logging), instrument observability (tracing, latency, usage tracking), integrate with LLM providers and the evaluation platform, and collaborate with research and product teams to scale and harden systems for bursty, unpredictable traffic.
Top Skills:
Anthropic ApiAWSDistributed TracingGCPGoGoogle Llm ApisKubernetesOpenai ApiPostgresRedisRustTerraform
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills:
C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
Edtech
Maintain and improve site performance, uptime, and scalability. Build monitoring, alerts, runbooks, and deployment tooling. Troubleshoot across the stack, partner with application teams, propose architecture changes, and support on-call operations to ensure reliable product delivery.
Top Skills:
AWSBashCC++DatabasesDockerGCPJavaKubernetesLoad BalancersMessage QueuesMonitoring DashboardsObservability ToolsPerlPythonWeb Servers
Cloud • Software
Lead incident detection and resolution, design automation and self-healing systems, build observability and AI/ML-powered operational tooling, define SLOs/SLIs and error budgets, mentor engineers, and operate in a 24/7 global SRE model to improve reliability and reduce toil.
Top Skills:
AirflowArgo WorkflowsCachingCi/CdClaude CodeCodexCursorDatadogDnsDockerElkGithub CopilotGoGrafanaHTTPKubernetesLinuxLoad BalancingPrometheusPythonSplunkTemporalUnix
Cloud • Information Technology • Internet of Things • Professional Services • Software
Operate and scale ThousandEyes Federal region infrastructure in a FedRAMP-compliant AWS environment. Design, deploy, and automate cloud-native services, implement IaC, monitor and audit systems, collaborate with security teams to remediate vulnerabilities, participate in 24x7 incident response and capacity planning, and ensure platform reliability, performance, and compliance.
Top Skills:
AWSFedrampGoKubernetesLinuxPuppetPythonTerraformUnixUs Govcloud
Artificial Intelligence • Legal Tech • Software • Generative AI
Lead and own the release and deployment process, manage GitHub workflows and Actions, build and maintain AWS infrastructure and observability, automate deployments and internal tooling, respond to incidents and be on-call, contribute code to reliability tooling, and support global/offshore teams across time zones.
Top Skills:
AWSBashChatgptCi/CdClaudeEc2GitGithub ActionsIamLambdaMetricsObservability (LogsPostgres SqlPythonRdsTraces)TypescriptVpc
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Site Reliability Engineer will ensure the reliability and performance of AI infrastructure, build core systems, handle incident response, and develop automation tools.
Top Skills:
AWSDatadogElkGCPGithub ActionsGitlab CiGoGrafanaJenkinsKubernetesLinuxPrometheusPulumiPythonRustTerraform
Social Media
Operate, scale, and improve a cloud-native platform on AWS and Kubernetes. Manage GitOps deployments with ArgoCD and Helm, provision infra with Terraform/Terragrunt, build CI/CD automation, enhance observability, respond to incidents, reduce operational toil through scripting, and collaborate with security and application teams to improve reliability and platform guardrails.
Top Skills:
ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonTerraformTerragrunt
Artificial Intelligence
The SRE/Infrastructure Engineer will manage Terraform and Kubernetes across cloud platforms, ensuring scalable infrastructure. Responsibilities include multi-cloud deployments, observability, and creating reusable components.
Top Skills:
AWSAzureCloudflareGCPKubernetesTerraform
Reposted 8 Days AgoSaved
Artificial Intelligence • Software
Design, build, and scale control- and data-plane infrastructure for distributed AI workloads. Improve reliability, performance, scheduling, and observability for Ray clusters across cloud and on-prem environments. Support accelerator integration, container image management, and provide on-call troubleshooting and cross-team collaboration.
Top Skills:
AWSAzureContainersGCPGoGpusGrafanaKubernetesLinuxPrometheusPythonRayTpusVms
Cloud
The role involves building and managing observability infrastructure in GCP, automating deployments, and optimizing data processes for high reliability.
Top Skills:
GkeGoGCPGrafanaKubernetesOpentelemetryPythonRubySplunkTerraform
Artificial Intelligence • Cloud • Information Technology • Software
As a Staff SRE, you will ensure the reliability and performance of Andromeda's GPU infrastructure, lead incident responses, build observability systems, and mentor engineers, while collaborating closely with engineering and customers.
Top Skills:
AnsibleCudaGoHelmKubernetesLinuxNcclNvidiaPythonRustSlurmTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Bay Area, CA Companies Hiring SRE Engineers
See AllPopular San Francisco Bay Area, CA Engineering Job Searches
Engineer Jobs in San Francisco Bay Area, CA
Software Engineer Jobs in San Francisco Bay Area, CA
Android Developer Jobs in San Francisco Bay Area, CA
C# Jobs in San Francisco Bay Area, CA
C++ Jobs in San Francisco Bay Area, CA
DevOps Jobs in San Francisco Bay Area, CA
Front End Developer Jobs in San Francisco Bay Area, CA
Golang Jobs in San Francisco Bay Area, CA
Hardware Engineer Jobs in San Francisco Bay Area, CA
iOS Developer Jobs in San Francisco Bay Area, CA
Java Developer Jobs in San Francisco Bay Area, CA
Javascript Jobs in San Francisco Bay Area, CA
Linux Jobs in San Francisco Bay Area, CA
Engineering Manager Jobs in San Francisco Bay Area, CA
.NET Developer Jobs in San Francisco Bay Area, CA
PHP Developer Jobs in San Francisco Bay Area, CA
Python Jobs in San Francisco Bay Area, CA
QA Jobs in San Francisco Bay Area, CA
Ruby Jobs in San Francisco Bay Area, CA
Salesforce Developer Jobs in San Francisco Bay Area, CA
Scala Jobs in San Francisco Bay Area, CA
Application Engineer Jobs in San Francisco, CA
Associate Software Engineer Jobs in San Francisco Bay Area, CA
Automation Engineer Jobs in San Francisco Bay Area, CA
AWS Engineer Jobs in San Francisco, CA
Backend Engineer Jobs in San Francisco, CA
Cloud Engineer Jobs in San Francisco Bay Area, CA
Controls Engineer Jobs in San Francisco Bay Area, CA
CTO Jobs in San Francisco Bay Area, CA
Design Engineer Jobs in San Francisco Bay Area, CA
DevOps Engineer Jobs in San Francisco Bay Area, CA
Director of Engineering Jobs in San Francisco, CA
Electrical Engineer Jobs in San Francisco Bay Area, CA
Embedded Software Engineer Jobs in San Francisco Bay Area, CA
Field Engineer Jobs in San Francisco Bay Area, CA
Firmware Engineer Jobs in San Francisco, CA
Full-Stack Engineer Jobs in San Francisco, CA
Game Engineer Jobs in San Francisco, CA
Infrastructure Engineer Jobs in San Francisco, CA
Manufacturing Engineer Jobs in San Francisco Bay Area, CA
Mechanical Design Engineer Jobs in San Francisco Bay Area, CA
Mechanical Engineer Jobs in San Francisco Bay Area, CA
Mechatronics Engineering Jobs in San Francisco Bay Area, CA
Network Engineer Jobs in San Francisco Bay Area, CA
Platform Engineer Jobs in San Francisco, CA
Process Engineer Jobs in San Francisco Bay Area, CA
Project Engineer Jobs in San Francisco Bay Area, CA
QA Engineer Jobs in San Francisco Bay Area, CA
Robotics Engineer Jobs in San Francisco Bay Area, CA
Security Engineer Jobs in San Francisco Bay Area, CA
Software Engineering Manager Jobs in San Francisco, CA
Software Test Engineer Jobs in San Francisco, CA
SRE Engineer Jobs in San Francisco Bay Area, CA
Systems Engineer Jobs in San Francisco Bay Area, CA
VP of Engineering Jobs in San Francisco, CA
All Filters
Total selected ()
No Results
No Results









.png)

























