Maximum of 25 job preferences reached.
Top Engineering Jobs in San Francisco, CA
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Build developer tooling, demos, and integrations for Runpod; produce technical content and run scalable programs (content, events, community); attend and speak at developer/AI conferences; surface product feedback and ship solutions autonomously to improve developer onboarding and platform adoption.
Top Skills:
Agent FrameworksAi Libraries And FrameworksGoGpu ComputeJavaScriptMl InfrastructurePythonSdks
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design, scale, and operate Runpod’s distributed storage platform across network volumes, local NVMe, and S3-compatible object storage. Tune storage and high-speed network performance, lead capacity expansions and migrations, write production automation and control-plane code, extend APIs, and build observability, dashboards, SLOs, and alerts. The role also participates in on-call operations, incident response, infrastructure-as-code practices, and long-term architecture decisions for petabyte-scale AI infrastructure.
Top Skills:
APIsCephCi/CdCsiDatadogEthernetGoGpfsGpudirect StorageGrafanaInfinibandInfrastructure As CodeIscsiKubernetesLinuxLustreMinioMoosefsNfsNvmeNvme-OfPrometheusPythonRdmaRoceRustS3SmbSpectrum ScaleStatefulsetsVastWekafsZfs
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design, build, and maintain cloud infrastructure software (primarily in Go) to manage VMs across global data centers. Collaborate on product requirements, participate in code reviews, optimize performance and reliability, contribute to architecture decisions, and stay current with industry trends.
Top Skills:
CDockerGoHypervisorJavaScriptKernel/Driver DebuggingLinuxLxcPythonPython Machine Learning LibrariesRustTypescriptVirtual Machines
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and scale core cloud and bare-metal infrastructure including SRE, global and HPC networking, and distributed storage. Define SLOs/SLAs, incident response, observability, and IaC. Architect InfiniBand/RoCE networks and high-performance storage to support large GPU workloads. Hire and mentor managers and senior ICs, partner with product and program teams to forecast capacity and drive reliability, throughput, and low-latency infrastructure at scale.
Top Skills:
AnsibleBare-MetalBgpCephContainer OrchestrationInfinibandInfrastructure As CodeKubernetesLustreNvlinkNvme-OfObservabilityRdmaRoceSpine-Leaf ArchitectureTerraformWeka
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead a product engineering team to own roadmaps, design scalable cloud-native systems, ship customer-facing features, ensure quality and reliability, hire and mentor engineers, and coordinate cross-functional delivery.
Top Skills:
APIsContainer RuntimesControl PlanesData StoresDockerEventingGoGpu/Accelerator WorkflowsKubernetesLinuxMicroservicesNetworkingOrchestrationPythonStorageTypescript
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
All Filters
Total selected ()
No Results
No Results


