Top Tech Jobs & Startup Jobs in San Francisco Bay Area, CA

One Month AgoSaved
Hybrid
2 Locations
236K-315K Annually
Expert/Leader
236K-315K Annually
Expert/Leader
Software
Lead SOX testing and ICFR for finance and operations, perform walkthroughs, design assessments, testing, and remediation. Manage documentation, coordinate with finance and cross-functional stakeholders, oversee co-sourced testers, and report results to senior management and the Audit Committee.
Top Skills: AuditboardErp SystemsFinancial Close ToolsJIRAWorkiva
One Month AgoSaved
Remote or Hybrid
2 Locations
226K-355K Annually
Senior level
226K-355K Annually
Senior level
Software
Lead technical pre-sales for Lambda's AI cloud: design and propose high-performance GPU infrastructure, run PoCs and benchmarks, architect and optimize distributed AI/ML workloads, advise on networking/storage, collaborate with sales and product teams, create enablement materials, and mentor junior SEs.
Top Skills: AnsibleC/C++ (Cuda)DockerGoHgxInfinibandKubernetesNemoNfsNvidia GpuNvlinkNvme-OfPythonPyTorchRoceSlurmTensorrt-LlmTerraformVastVllmWeka
Reposted One Month AgoSaved
Remote or Hybrid
2 Locations
240K-312K Annually
Senior level
240K-312K Annually
Senior level
Software
Operate and scale Lambda's multi-tenant cloud networking and SDN infrastructure; run Kubernetes control plane and SmartNIC dataplane software; build automation, CI/CD and GitOps workflows; deploy monitoring and observability; collaborate across teams, drive incident response and on-call rotation, capacity planning, and postmortems to improve reliability.
Top Skills: AnsibleCCi/CdDpdkGitopsGoHelmKubernetesLinuxMonitoring/ObservabilityOpenstack NeutronOvnOvsPythonSmartnicsSr-IovTerraform
Reposted 25 Days AgoSaved
In-Office or Remote
USA
125K-195K Annually
Senior level
125K-195K Annually
Senior level
Software
The Senior Incident Manager leads incident response for AI infrastructure, coordinates cross-team collaboration, oversees incident resolution and post-incident reviews to enhance operational resilience.
Top Skills: Cloud PlatformsDatadogGpu ClustersGrafanaJIRANetworkingPagerdutyPrometheusServicenow
27 Days AgoSaved
Remote
USA
214K-285K Annually
Mid level
214K-285K Annually
Mid level
Software
Owns technical customer deployments from signed contract through production readiness. Validates GPU, cloud, connectivity, storage, and compute configurations; troubleshoots issues; coordinates Infrastructure, Engineering, Product, and Data Center teams; tracks risks and dependencies; provides stakeholder updates; guides onboarding; and develops reusable runbooks and automation before transitioning stable customers to the account team.
Top Skills: Cloud PlatformsCompute InfrastructureGpuHpc InfrastructureKubernetesLinuxNetworkingStorage
One Month AgoSaved
Remote
USA
122K-162K Annually
Mid level
122K-162K Annually
Mid level
Software
Provide senior-level technical escalation and support for GPU/HPC cloud infrastructure. Troubleshoot hardware, drivers, kernel, networking, and workload issues; perform root-cause analysis across clusters; build automations and documentation; mentor junior engineers; participate in on-call rotations and incident response; collaborate with engineering to implement permanent fixes.
Top Skills: AnsibleCi/CdCudaDatadogDockerFirewallsGpudirect RdmaGrafanaInfinibandKubernetesLinuxNcclNvidia GpusNvlinkPrometheusRoceSlurmTcp/IpTerraformVpn
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account