Top Infrastructure Engineer Jobs in San Francisco, CA

14 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
140K-274K Annually
Senior level
140K-274K Annually
Senior level
Artificial Intelligence • Software • Generative AI
Build and operate highly available, scalable infrastructure for WRITER’s enterprise AI platform. Responsibilities include SRE, DevOps, platform engineering, cloud infrastructure, Kubernetes, Terraform, automation, observability, incident response, post-mortems, SLOs, on-call operations, and reliability improvements. The role requires cross-functional collaboration, end-to-end ownership, systemic problem-solving, and daily use of AI-assisted development and operational tooling.
Top Skills: AWSAzureClaude CodeCodexDroidElkGCPGoGrafanaHelmKubernetesPrometheusPulumiPythonTerraform
15 Days AgoSaved
In-Office
San Francisco Bay Area, CA
150K-250K Annually
Mid level
150K-250K Annually
Mid level
Artificial Intelligence • Robotics • Automation • Manufacturing
Operate and improve infrastructure spanning AWS, Kubernetes, on-premises and edge systems, deployment pipelines, GPU workloads, networking, and observability. Build infrastructure as code with Terraform and GitOps, troubleshoot networking issues, support factory-floor devices, maintain trustworthy alerts, and apply secure access and secrets-management practices. The role may also involve Kubernetes on-premises, arm64 edge fleets, over-the-air updates, WireGuard networks, robotics hardware, and ITAR or CMMC environments.
Top Skills: Arm64AWSDnsGitopsGoInfrastructure Ci/CdKubernetesLinuxObservabilityOverlay NetworksPythonTerraformTlsVpcVpnWireguard
15 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
100K-200K Annually
Mid level
100K-200K Annually
Mid level
Software • Industrial • Generative AI
Build and improve AI agents that generate trusted engineering deliverables for physical infrastructure projects. Apply expertise in process, electrical, mechanical, civil, I&C, piping, structural, or environmental engineering to define correct outputs, collaborate with software teams and Fortune 500 clients, and own deliverables across data center, energy, refinery, and chemical plant projects. Travel to partner sites approximately 10% of the time.
Top Skills: Claude CodeCodex
11 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Reposted One Month AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
165K-165K Annually
Senior level
165K-165K Annually
Senior level
Security • Cybersecurity
Design, build, and maintain scalable AWS and Azure infrastructure using Infrastructure-as-Code. Develop automation, deployment pipelines, and production release processes. Contribute code to cloud services, manage networking and IAM, ensure availability and security, and respond to incidents while collaborating on observability and security practices.
Top Skills: AnsibleAWSAzureDockerGoIamKubernetesPackerPythonTerraform
16 Days AgoSaved
In-Office
San Francisco Bay Area, CA
Internship
Internship
Artificial Intelligence • Machine Learning • Software • Infrastructure as a Service (IaaS)
Paid, full-time infrastructure engineering internship focused on distributed systems, persistent compute, storage, scheduling, orchestration, virtualization, networking, reliability, observability, and production AI-agent infrastructure. Interns will own and ship meaningful projects, write production-quality code and tests, investigate failures, and communicate architectural tradeoffs. The role is approximately three months, fully in person in San Francisco, with potential relocation and visa support depending on circumstances.
Top Skills: CC++Cloud InfrastructureContainersDistributed StorageFirecrackerGoHypervisorsKubernetesNetworkingObservabilityOperating SystemsRustVirtualization
17 Days AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
165K-200K Annually
Entry level
165K-200K Annually
Entry level
Artificial Intelligence • Computer Vision • Software
Build, secure, scale, and maintain cloud infrastructure supporting SaaS products, databases, storage, search, microservices, and machine-learning pipelines. Operate Kubernetes workloads, automate infrastructure with Terraform and Helm, improve reliability and observability, optimize costs, define SLOs and SLAs, participate in incident response and on-call rotations, address vulnerabilities, support customer security integrations, and maintain compliance with SOC 2, HIPAA, and GDPR requirements.
Top Skills: AWSBashCi/CdDockerGCPGithub ActionsGpusHelmInfrastructure-As-CodeJavaScriptKubernetesLlmsNode.jsPythonPyTorchSpaceliftTensorFlowTerraform
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
Expert/Leader
Expert/Leader
Aerospace
Lead cloud infrastructure, DevOps, CI/CD, build tooling, release management, testing pipelines, and data infrastructure for aerospace software teams. Develop reproducible workflows spanning firmware, spacecraft software, ground systems, and cloud services. Manage versioned datasets, ML training infrastructure, cloud reliability, security, and cost. Partner with engineering and mission operations teams, resolve complex system issues, contribute to architecture, and mentor engineers. The role may be onsite in Austin or Littleton, or remote with regular travel to Austin.
Top Skills: Automated TestingAWSAzureBazelBuild SystemsC++Ci/CdCloud InfrastructureCmakeData PipelinesDevOpsGCPHardware-In-The-Loop TestingInfrastructure As CodeLinuxMachine Learning InfrastructurePythonRelease ManagementTerraform
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
Senior level
Senior level
Software
Deploy, integrate, operate, and optimize high-performance NFS storage for GPU-accelerated Kubernetes and AI platforms. Configure Linux storage and networking, integrate CSI storage into k0s and K0rdent clusters, automate provisioning through Terraform/OpenTofu and GitOps, and build observability for capacity, performance, and reliability. Diagnose end-to-end storage issues across hybrid, edge, and air-gapped environments while establishing operational standards.
Top Skills: ArgocdCapiCluster ApiCsiDell PowerscaleFluxGitopsHarborK0RdentK0SKubernetesLinuxNfsOpentofuPkiRdmaTerraformTlsVast
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
110K-117K Annually
Senior level
110K-117K Annually
Senior level
Healthtech
Designs, builds, and operates secure, scalable AWS and Azure cloud infrastructure across multi-account environments. Develops Infrastructure as Code, CI/CD pipelines, automation, observability, and IAM governance. Supports production reliability, incident response, FinOps, FedRAMP and HIPAA compliance, and AI-assisted engineering practices. Partners with product, security, and engineering teams to establish cloud standards, guardrails, and reference architectures.
Top Skills: AnsibleAWSAws CdkAzureBashChefClaude CodeCloudFormationCloudwatchCursorEc2EcsEksGithub ActionsGithub CopilotIamJenkinsKubernetesPowershellPuppetPythonRdsS3SaltSaltstackTerraformVpc
8 Days AgoSaved
Remote
San Francisco Bay Area, CA
110K-118K Annually
Senior level
110K-118K Annually
Senior level
Healthtech
Owns the design, operation, resilience, security, and migration of HealthEdge’s hybrid cloud infrastructure across AWS, Azure, GCP, and on-premises environments. Responsibilities include Infrastructure as Code, disaster recovery, Kubernetes/EKS, IAM, server administration, CI/CD automation, monitoring, incident response, vulnerability management, compliance controls, cost governance, and technical documentation. The role participates in on-call support, leads infrastructure incident resolution, and guides teams on migration and recovery design.
Top Skills: Active DirectoryAmazon EksAWSAws CdkAws CloudformationAzureBashCi/CdCis BenchmarksDisa StigDnsFedrampFinopsGCPGroup PolicyHipaaIamInfrastructure As CodeKubernetesLinuxPowershellPythonSoc 2TerraformVpcWindows Server
Reposted One Month AgoSaved
Remote
San Francisco Bay Area, CA
145K-200K Annually
Senior level
145K-200K Annually
Senior level
Fintech • Information Technology • Software
The Senior Infrastructure Engineer will improve system reliability and efficiency, develop standards and tooling, and collaborate with engineering teams to optimize workflows and cloud infrastructure.
Top Skills: Aurora PostgresqlAWSCicdDatadogDockerKubernetesLinuxOpensearchPrometheusPythonSumologicTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
9 Days AgoSaved
Remote
San Francisco Bay Area, CA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Computer Vision • Machine Learning • Analytics
Design and scale infrastructure for cloud and on-premises deployments supporting enterprise AI workloads. Build containerized systems, Kubernetes and Docker Compose deployments, Helm charts, CI/CD pipelines, infrastructure automation, monitoring, and security solutions. Support customer production deployments, troubleshoot distributed systems, improve developer productivity, mentor engineers, and establish infrastructure standards across engineering teams.
Top Skills: AnsibleAWSAzureBashCi/CdDockerDocker ComposeGCPGithub ActionsGkeGoogle Cloud BuildHelmIamKubernetesMongoDBNoSQLPythonTerraformVpc
10 Days AgoSaved
Remote
San Francisco Bay Area, CA
150K-210K Annually
Senior level
150K-210K Annually
Senior level
Artificial Intelligence • Cybersecurity
Owns the design, operation, and improvement of infrastructure, build and release pipelines, CI/CD, and deployment automation. The role uses Terraform, Ansible, Helm, Go, Kubernetes, and DevSecOps practices across cloud, on-premises, air-gapped, and FedRAMP environments. Responsibilities include observability, secure endpoint exposure, runtime troubleshooting, software packaging, security-tool integration, and collaboration with Product and developers to deliver secure commercial software.
Top Skills: AnsibleApi SecurityAWSAzureCdnsCi/CdCmmcContainersCorsDastDevsecopsEventsFedrampGCPGitGoHelmInfrastructure As CodeKubernetesLoad BalancersLogsMetricsObservabilityPackerSastScaSelf-Hosted LlmsTerraformTracesZarf
25 Days AgoSaved
In-Office
San Francisco Bay Area, CA
160K-300K Annually
Senior level
160K-300K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Financial Services • Generative AI
Own Hebbia’s AWS infrastructure and developer platform end to end. Design multi-account architecture, networking, IAM, container orchestration, and infrastructure as code. Build CI/CD pipelines, preview environments, automation, observability, and cost controls. Strengthen security, compliance, secrets management, and audit readiness while preparing infrastructure for enterprise scale and multi-region deployments. Partner with product engineers to create reliable, secure, and efficient development and production systems.
Top Skills: AWSCi/CdCloud NetworkingContainer OrchestrationContainerizationGoIamInfrastructure As CodeMonitoringObservabilityPythonSecrets Management
20 Days AgoSaved
In-Office
San Francisco Bay Area, CA
159K-194K Annually
Entry level
159K-194K Annually
Entry level
Software
Develop and maintain SiFive’s verification platform infrastructure, libraries, tools, and methodologies for design verification. Build functionality using Chisel and CIRCT, create Scala-based verification libraries and generators, and develop Python tools for querying design intent and metadata. Collaborate with platform technology teams, support end-to-end verification flows, and educate teams on best practices. The role requires compiled and scripting language experience, modern software engineering knowledge, and experience deploying production software.
Top Skills: ChiselCirctEda ToolsMlirPythonScala
Reposted 25 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
150K-200K Annually
Mid level
150K-200K Annually
Mid level
Artificial Intelligence • Productivity • Software
Build and scale backend infrastructure for Notions database (Collections) to improve P95 performance, reliability, and concurrency. Work on distributed systems, microservices, high-volume writes, observability, and platform features; partner with product and senior engineers and use AI tools to accelerate debugging and testing.
Top Skills: DatabasesDistributed SystemsJavaScriptLoggingMicroservicesObservabilityTracingTypescript
11 Days AgoSaved
Remote
San Francisco Bay Area, CA
121K-157K Annually
Senior level
121K-157K Annually
Senior level
Healthtech
Own and improve Azure infrastructure while leading complex integrations, automation workflows, and cross-system implementations. Provide advanced escalation support for service desk issues, including SharePoint; diagnose interconnected platform problems; strengthen security, access controls, and system reliability; and maintain technical documentation. The role also supports internally developed applications, uses AI-assisted tools, and mentors engineers on engineering standards, testing, version control, and change management.
Top Skills: APIsGenesysIamInfrastructure As CodeAzureSharepoint
17 Days AgoSaved
Easy Apply
Remote
San Francisco Bay Area, CA
Easy Apply
218K-257K Annually
Senior level
218K-257K Annually
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Own and lead Coinbase’s developer infrastructure platforms, including CI, build systems, and deployment orchestration. Set technical direction for complex cross-team initiatives, improve build and deployment performance, manage reliability and observability, and drive platform migrations and adoption. Partner with engineering leaders, operate production systems, mentor engineers, and deliver scalable tools used across the organization.
Top Skills: Build SystemsCi/CdDeployment OrchestrationDistributed SystemsGenerative AiGoObservability
Reposted One Month AgoSaved
In-Office or Remote
San Francisco Bay Area, CA
190K-206K Annually
Senior level
190K-206K Annually
Senior level
Artificial Intelligence • Software
As a Senior Infrastructure Engineer, you will design and build core infrastructure, manage IaC, enhance system reliability, ensure security compliance, and mentor engineers.
Top Skills: AWSCloudwatchDatadogGrafanaGraphQLNew RelicNoSQLPrometheusSQLTerraform
23 Days AgoSaved
Hybrid
San Francisco Bay Area, CA
200K-275K Annually
Expert/Leader
200K-275K Annually
Expert/Leader
Healthtech
Own Onos Health’s production infrastructure, including AWS architecture, multi-region disaster recovery, backups, monitoring, SLOs, incident response, SOC 2 and HIPAA controls, Terraform, CI/CD, preview environments, and AI-agent deployment guardrails. Establish platform strategy, reliability practices, security automation, audit evidence, and on-call operations as the company’s first dedicated infrastructure hire. The role also requires technical leadership, delegation, and communication with non-technical stakeholders.
Top Skills: Amazon BedrockAmazon EcsAmazon RdsAmazon S3AWSAws GlueAws KmsCeleryCelery BeatClaude CodeClickhouseCoderabbitaiCodexDjangoDjango-NinjaDjango-TenantsDrataGitHipaaHitrustIamIso 27001KyvernoLangfuseLinearNext.JsNuqsOpaPostgresPythonShadcn/UiSoc 2Tanstack QueryTerraformTypescriptVantaZod
Reposted YesterdaySaved
In-Office or Remote
San Francisco Bay Area, CA
296K-444K Annually
Senior level
296K-444K Annually
Senior level
Consumer Web • Mobile
Lead Patreon’s Infrastructure team, owning its roadmap, execution, reliability, observability, incident response, on-call structure, SLAs/SLOs, and production readiness. Manage and grow engineers, including hiring, performance, and career development. Partner with Product Engineering, Security, and Data leaders to improve developer productivity and platform scalability. The role requires extensive engineering and management experience, infrastructure or platform leadership, cloud and Kubernetes expertise, and experience operating highly available production systems.
Top Skills: Ai-Assisted Software Development ToolsAWSCi/CdGCPInfrastructure As CodeKubernetes
YesterdaySaved
In-Office
San Francisco Bay Area, CA
208K-260K Annually
Senior level
208K-260K Annually
Senior level
Edtech • Enterprise Web • HR Tech • Software
Build and operate backend services, APIs, shared platform capabilities, and developer tooling for Handshake AI and internal AI initiatives. Responsibilities include service architecture, database design, performance, scalability, reliability, observability, and developer experience. The role contributes to LLM and agentic systems, including orchestration, evaluation, integrations, and failure handling. It also involves architectural decisions, incident response, cross-team collaboration, technical planning, code reviews, and mentoring engineers.
Top Skills: Agent OrchestrationAi EvaluationDistributed SystemsDockerElasticsearchEvent-Driven ArchitectureGCPGraphQLKubernetesLlmsNext.JsObservabilityPostgresPrompt ManagementRagRedisTemporalTerraformTypescript
YesterdaySaved
Remote or Hybrid
San Francisco Bay Area, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
Top Skills: APIsBare-Metal InfrastructureBmcContainerizationDell HardwareGpu ServersHpcIpmiJuniper NetworksLinuxOrchestrationPalo Alto FirewallsPxe/IpxePythonRedfishSonicVast
YesterdaySaved
Remote or Hybrid
San Francisco Bay Area, CA
160K-200K Annually
Senior level
160K-200K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Operate and scale large GPU infrastructure platforms, including Linux systems, bare-metal environments, provisioning workflows, observability, and reliability automation. Responsibilities include platform deployment, incident response, break/fix operations, customer provisioning, on-call participation, infrastructure troubleshooting, and collaboration across engineering, networking, customer success, and software teams. The role also develops automation to reduce manual work and improve operational efficiency.
Top Skills: AnsibleAWSBashCephDellElk StackEthernetGitopsGoGpuInfinibandJuniperKubernetesLinuxNfsPalo AltoPrometheusPythonSonicTerraformUbuntuVast
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account