Maximum of 25 job preferences reached.
Top Infrastructure Engineer Jobs in San Francisco, CA
Artificial Intelligence • Software • Generative AI
Build and operate highly available, scalable infrastructure for WRITER’s enterprise AI platform. Responsibilities include SRE, DevOps, platform engineering, cloud infrastructure, Kubernetes, Terraform, automation, observability, incident response, post-mortems, SLOs, on-call operations, and reliability improvements. The role requires cross-functional collaboration, end-to-end ownership, systemic problem-solving, and daily use of AI-assisted development and operational tooling.
Top Skills:
AWSAzureClaude CodeCodexDroidElkGCPGoGrafanaHelmKubernetesPrometheusPulumiPythonTerraform
Artificial Intelligence • Robotics • Automation • Manufacturing
Operate and improve infrastructure spanning AWS, Kubernetes, on-premises and edge systems, deployment pipelines, GPU workloads, networking, and observability. Build infrastructure as code with Terraform and GitOps, troubleshoot networking issues, support factory-floor devices, maintain trustworthy alerts, and apply secure access and secrets-management practices. The role may also involve Kubernetes on-premises, arm64 edge fleets, over-the-air updates, WireGuard networks, robotics hardware, and ITAR or CMMC environments.
Top Skills:
Arm64AWSDnsGitopsGoInfrastructure Ci/CdKubernetesLinuxObservabilityOverlay NetworksPythonTerraformTlsVpcVpnWireguard
Software • Industrial • Generative AI
Build and improve AI agents that generate trusted engineering deliverables for physical infrastructure projects. Apply expertise in process, electrical, mechanical, civil, I&C, piping, structural, or environmental engineering to define correct outputs, collaborate with software teams and Fortune 500 clients, and own deliverables across data center, energy, refinery, and chemical plant projects. Travel to partner sites approximately 10% of the time.
Top Skills:
Claude CodeCodex
11 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Security • Cybersecurity
Design, build, and maintain scalable AWS and Azure infrastructure using Infrastructure-as-Code. Develop automation, deployment pipelines, and production release processes. Contribute code to cloud services, manage networking and IAM, ensure availability and security, and respond to incidents while collaborating on observability and security practices.
Top Skills:
AnsibleAWSAzureDockerGoIamKubernetesPackerPythonTerraform
Artificial Intelligence • Machine Learning • Software • Infrastructure as a Service (IaaS)
Paid, full-time infrastructure engineering internship focused on distributed systems, persistent compute, storage, scheduling, orchestration, virtualization, networking, reliability, observability, and production AI-agent infrastructure. Interns will own and ship meaningful projects, write production-quality code and tests, investigate failures, and communicate architectural tradeoffs. The role is approximately three months, fully in person in San Francisco, with potential relocation and visa support depending on circumstances.
Top Skills:
CC++Cloud InfrastructureContainersDistributed StorageFirecrackerGoHypervisorsKubernetesNetworkingObservabilityOperating SystemsRustVirtualization
Artificial Intelligence • Computer Vision • Software
Build, secure, scale, and maintain cloud infrastructure supporting SaaS products, databases, storage, search, microservices, and machine-learning pipelines. Operate Kubernetes workloads, automate infrastructure with Terraform and Helm, improve reliability and observability, optimize costs, define SLOs and SLAs, participate in incident response and on-call rotations, address vulnerabilities, support customer security integrations, and maintain compliance with SOC 2, HIPAA, and GDPR requirements.
Top Skills:
AWSBashCi/CdDockerGCPGithub ActionsGpusHelmInfrastructure-As-CodeJavaScriptKubernetesLlmsNode.jsPythonPyTorchSpaceliftTensorFlowTerraform
Aerospace
Lead cloud infrastructure, DevOps, CI/CD, build tooling, release management, testing pipelines, and data infrastructure for aerospace software teams. Develop reproducible workflows spanning firmware, spacecraft software, ground systems, and cloud services. Manage versioned datasets, ML training infrastructure, cloud reliability, security, and cost. Partner with engineering and mission operations teams, resolve complex system issues, contribute to architecture, and mentor engineers. The role may be onsite in Austin or Littleton, or remote with regular travel to Austin.
Top Skills:
Automated TestingAWSAzureBazelBuild SystemsC++Ci/CdCloud InfrastructureCmakeData PipelinesDevOpsGCPHardware-In-The-Loop TestingInfrastructure As CodeLinuxMachine Learning InfrastructurePythonRelease ManagementTerraform
Software
Deploy, integrate, operate, and optimize high-performance NFS storage for GPU-accelerated Kubernetes and AI platforms. Configure Linux storage and networking, integrate CSI storage into k0s and K0rdent clusters, automate provisioning through Terraform/OpenTofu and GitOps, and build observability for capacity, performance, and reliability. Diagnose end-to-end storage issues across hybrid, edge, and air-gapped environments while establishing operational standards.
Top Skills:
ArgocdCapiCluster ApiCsiDell PowerscaleFluxGitopsHarborK0RdentK0SKubernetesLinuxNfsOpentofuPkiRdmaTerraformTlsVast
Healthtech
Designs, builds, and operates secure, scalable AWS and Azure cloud infrastructure across multi-account environments. Develops Infrastructure as Code, CI/CD pipelines, automation, observability, and IAM governance. Supports production reliability, incident response, FinOps, FedRAMP and HIPAA compliance, and AI-assisted engineering practices. Partners with product, security, and engineering teams to establish cloud standards, guardrails, and reference architectures.
Top Skills:
AnsibleAWSAws CdkAzureBashChefClaude CodeCloudFormationCloudwatchCursorEc2EcsEksGithub ActionsGithub CopilotIamJenkinsKubernetesPowershellPuppetPythonRdsS3SaltSaltstackTerraformVpc
Healthtech
Owns the design, operation, resilience, security, and migration of HealthEdge’s hybrid cloud infrastructure across AWS, Azure, GCP, and on-premises environments. Responsibilities include Infrastructure as Code, disaster recovery, Kubernetes/EKS, IAM, server administration, CI/CD automation, monitoring, incident response, vulnerability management, compliance controls, cost governance, and technical documentation. The role participates in on-call support, leads infrastructure incident resolution, and guides teams on migration and recovery design.
Top Skills:
Active DirectoryAmazon EksAWSAws CdkAws CloudformationAzureBashCi/CdCis BenchmarksDisa StigDnsFedrampFinopsGCPGroup PolicyHipaaIamInfrastructure As CodeKubernetesLinuxPowershellPythonSoc 2TerraformVpcWindows Server
Fintech • Information Technology • Software
The Senior Infrastructure Engineer will improve system reliability and efficiency, develop standards and tooling, and collaborate with engineering teams to optimize workflows and cloud infrastructure.
Top Skills:
Aurora PostgresqlAWSCicdDatadogDockerKubernetesLinuxOpensearchPrometheusPythonSumologicTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Computer Vision • Machine Learning • Analytics
Design and scale infrastructure for cloud and on-premises deployments supporting enterprise AI workloads. Build containerized systems, Kubernetes and Docker Compose deployments, Helm charts, CI/CD pipelines, infrastructure automation, monitoring, and security solutions. Support customer production deployments, troubleshoot distributed systems, improve developer productivity, mentor engineers, and establish infrastructure standards across engineering teams.
Top Skills:
AnsibleAWSAzureBashCi/CdDockerDocker ComposeGCPGithub ActionsGkeGoogle Cloud BuildHelmIamKubernetesMongoDBNoSQLPythonTerraformVpc
Artificial Intelligence • Cybersecurity
Owns the design, operation, and improvement of infrastructure, build and release pipelines, CI/CD, and deployment automation. The role uses Terraform, Ansible, Helm, Go, Kubernetes, and DevSecOps practices across cloud, on-premises, air-gapped, and FedRAMP environments. Responsibilities include observability, secure endpoint exposure, runtime troubleshooting, software packaging, security-tool integration, and collaboration with Product and developers to deliver secure commercial software.
Top Skills:
AnsibleApi SecurityAWSAzureCdnsCi/CdCmmcContainersCorsDastDevsecopsEventsFedrampGCPGitGoHelmInfrastructure As CodeKubernetesLoad BalancersLogsMetricsObservabilityPackerSastScaSelf-Hosted LlmsTerraformTracesZarf
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Financial Services • Generative AI
Own Hebbia’s AWS infrastructure and developer platform end to end. Design multi-account architecture, networking, IAM, container orchestration, and infrastructure as code. Build CI/CD pipelines, preview environments, automation, observability, and cost controls. Strengthen security, compliance, secrets management, and audit readiness while preparing infrastructure for enterprise scale and multi-region deployments. Partner with product engineers to create reliable, secure, and efficient development and production systems.
Top Skills:
AWSCi/CdCloud NetworkingContainer OrchestrationContainerizationGoIamInfrastructure As CodeMonitoringObservabilityPythonSecrets Management
Software
Develop and maintain SiFive’s verification platform infrastructure, libraries, tools, and methodologies for design verification. Build functionality using Chisel and CIRCT, create Scala-based verification libraries and generators, and develop Python tools for querying design intent and metadata. Collaborate with platform technology teams, support end-to-end verification flows, and educate teams on best practices. The role requires compiled and scripting language experience, modern software engineering knowledge, and experience deploying production software.
Top Skills:
ChiselCirctEda ToolsMlirPythonScala
Artificial Intelligence • Productivity • Software
Build and scale backend infrastructure for Notions database (Collections) to improve P95 performance, reliability, and concurrency. Work on distributed systems, microservices, high-volume writes, observability, and platform features; partner with product and senior engineers and use AI tools to accelerate debugging and testing.
Top Skills:
DatabasesDistributed SystemsJavaScriptLoggingMicroservicesObservabilityTracingTypescript
Healthtech
Own and improve Azure infrastructure while leading complex integrations, automation workflows, and cross-system implementations. Provide advanced escalation support for service desk issues, including SharePoint; diagnose interconnected platform problems; strengthen security, access controls, and system reliability; and maintain technical documentation. The role also supports internally developed applications, uses AI-assisted tools, and mentors engineers on engineering standards, testing, version control, and change management.
Top Skills:
APIsGenesysIamInfrastructure As CodeAzureSharepoint
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Own and lead Coinbase’s developer infrastructure platforms, including CI, build systems, and deployment orchestration. Set technical direction for complex cross-team initiatives, improve build and deployment performance, manage reliability and observability, and drive platform migrations and adoption. Partner with engineering leaders, operate production systems, mentor engineers, and deliver scalable tools used across the organization.
Top Skills:
Build SystemsCi/CdDeployment OrchestrationDistributed SystemsGenerative AiGoObservability
Artificial Intelligence • Software
As a Senior Infrastructure Engineer, you will design and build core infrastructure, manage IaC, enhance system reliability, ensure security compliance, and mentor engineers.
Top Skills:
AWSCloudwatchDatadogGrafanaGraphQLNew RelicNoSQLPrometheusSQLTerraform
Healthtech
Own Onos Health’s production infrastructure, including AWS architecture, multi-region disaster recovery, backups, monitoring, SLOs, incident response, SOC 2 and HIPAA controls, Terraform, CI/CD, preview environments, and AI-agent deployment guardrails. Establish platform strategy, reliability practices, security automation, audit evidence, and on-call operations as the company’s first dedicated infrastructure hire. The role also requires technical leadership, delegation, and communication with non-technical stakeholders.
Top Skills:
Amazon BedrockAmazon EcsAmazon RdsAmazon S3AWSAws GlueAws KmsCeleryCelery BeatClaude CodeClickhouseCoderabbitaiCodexDjangoDjango-NinjaDjango-TenantsDrataGitHipaaHitrustIamIso 27001KyvernoLangfuseLinearNext.JsNuqsOpaPostgresPythonShadcn/UiSoc 2Tanstack QueryTerraformTypescriptVantaZod
Consumer Web • Mobile
Lead Patreon’s Infrastructure team, owning its roadmap, execution, reliability, observability, incident response, on-call structure, SLAs/SLOs, and production readiness. Manage and grow engineers, including hiring, performance, and career development. Partner with Product Engineering, Security, and Data leaders to improve developer productivity and platform scalability. The role requires extensive engineering and management experience, infrastructure or platform leadership, cloud and Kubernetes expertise, and experience operating highly available production systems.
Top Skills:
Ai-Assisted Software Development ToolsAWSCi/CdGCPInfrastructure As CodeKubernetes
Edtech • Enterprise Web • HR Tech • Software
Build and operate backend services, APIs, shared platform capabilities, and developer tooling for Handshake AI and internal AI initiatives. Responsibilities include service architecture, database design, performance, scalability, reliability, observability, and developer experience. The role contributes to LLM and agentic systems, including orchestration, evaluation, integrations, and failure handling. It also involves architectural decisions, incident response, cross-team collaboration, technical planning, code reviews, and mentoring engineers.
Top Skills:
Agent OrchestrationAi EvaluationDistributed SystemsDockerElasticsearchEvent-Driven ArchitectureGCPGraphQLKubernetesLlmsNext.JsObservabilityPostgresPrompt ManagementRagRedisTemporalTerraformTypescript
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
Top Skills:
APIsBare-Metal InfrastructureBmcContainerizationDell HardwareGpu ServersHpcIpmiJuniper NetworksLinuxOrchestrationPalo Alto FirewallsPxe/IpxePythonRedfishSonicVast
Artificial Intelligence • Machine Learning • Software
Operate and scale large GPU infrastructure platforms, including Linux systems, bare-metal environments, provisioning workflows, observability, and reliability automation. Responsibilities include platform deployment, incident response, break/fix operations, customer provisioning, on-call participation, infrastructure troubleshooting, and collaboration across engineering, networking, customer success, and software teams. The role also develops automation to reduce manual work and improve operational efficiency.
Top Skills:
AnsibleAWSBashCephDellElk StackEthernetGitopsGoGpuInfinibandJuniperKubernetesLinuxNfsPalo AltoPrometheusPythonSonicTerraformUbuntuVast
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top San Francisco Bay Area, CA Companies Hiring Infrastructure Engineers
See AllPopular San Francisco Bay Area, CA Engineering Job Searches
Engineering Jobs in San Francisco Bay Area, CA
Software Engineer Jobs in San Francisco Bay Area, CA
Android Developer Jobs in San Francisco Bay Area, CA
C# Jobs in San Francisco Bay Area, CA
C++ Jobs in San Francisco Bay Area, CA
DevOps Jobs in San Francisco Bay Area, CA
Front End Developer Jobs in San Francisco Bay Area, CA
Golang Jobs in San Francisco Bay Area, CA
Hardware Engineer Jobs in San Francisco Bay Area, CA
iOS Developer Jobs in San Francisco Bay Area, CA
Java Developer Jobs in San Francisco Bay Area, CA
Javascript Jobs in San Francisco Bay Area, CA
Linux Jobs in San Francisco Bay Area, CA
Engineering Manager Jobs in San Francisco Bay Area, CA
.NET Developer Jobs in San Francisco Bay Area, CA
PHP Developer Jobs in San Francisco Bay Area, CA
Python Jobs in San Francisco Bay Area, CA
QA Jobs in San Francisco Bay Area, CA
Ruby Jobs in San Francisco Bay Area, CA
Salesforce Developer Jobs in San Francisco Bay Area, CA
Scala Jobs in San Francisco Bay Area, CA
Application Engineer Jobs in San Francisco, CA
Associate Software Engineer Jobs in San Francisco Bay Area, CA
Automation Engineer Jobs in San Francisco Bay Area, CA
AWS Engineer Jobs in San Francisco, CA
Backend Engineer Jobs in San Francisco, CA
Cloud Engineer Jobs in San Francisco Bay Area, CA
Controls Engineer Jobs in San Francisco Bay Area, CA
CTO Jobs in San Francisco Bay Area, CA
Design Engineer Jobs in San Francisco Bay Area, CA
DevOps Engineer Jobs in San Francisco Bay Area, CA
Director of Engineering Jobs in San Francisco, CA
Electrical Engineering Jobs in San Francisco Bay Area, CA
Embedded Software Engineer Jobs in San Francisco Bay Area, CA
Field Engineer Jobs in San Francisco Bay Area, CA
Firmware Engineer Jobs in San Francisco, CA
Full-Stack Engineer Jobs in San Francisco, CA
Game Engineer Jobs in San Francisco, CA
Infrastructure Engineer Jobs in San Francisco, CA
Manufacturing Engineer Jobs in San Francisco Bay Area, CA
Mechanical Design Engineer Jobs in San Francisco Bay Area, CA
Mechanical Engineering Jobs in San Francisco Bay Area, CA
Mechatronics Engineering Jobs in San Francisco Bay Area, CA
Network Engineer Jobs in San Francisco Bay Area, CA
Platform Engineer Jobs in San Francisco, CA
Process Engineer Jobs in San Francisco Bay Area, CA
Project Engineer Jobs in San Francisco Bay Area, CA
QA Engineer Jobs in San Francisco Bay Area, CA
Robotics Engineer Jobs in San Francisco Bay Area, CA
Security Engineer Jobs in San Francisco Bay Area, CA
Software Engineering Manager Jobs in San Francisco, CA
Software Test Engineer Jobs in San Francisco, CA
SRE Engineer Jobs in San Francisco Bay Area, CA
Systems Engineer Jobs in San Francisco Bay Area, CA
VP of Engineering Jobs in San Francisco, CA
All Filters
Total selected ()
No Results
No Results
















.png)















