Top Tech Jobs & Startup Jobs in San Francisco Bay Area, CA

23 Days AgoSaved
Remote or Hybrid
3 Locations
235K-315K Annually
Expert/Leader
235K-315K Annually
Expert/Leader
Software
Lead and scale Lambda Cloud's enterprise sales organization: recruit and coach sellers, define GTM and commercial terms, manage pipeline and forecasting, align capacity allocation with customer demand, close strategic, high-value accounts, and channel customer feedback into product and capacity roadmaps.
Top Skills: Aws MarketplaceAzure MarketplaceGcp MarketplaceGpuGpu ComputeHpcNvidia
25 Days AgoSaved
Remote or Hybrid
2 Locations
240K-312K Annually
Senior level
240K-312K Annually
Senior level
Software
Operate and scale Lambda's multi-tenant cloud networking and SDN infrastructure; run Kubernetes control plane and SmartNIC dataplane software; build automation, CI/CD and GitOps workflows; deploy monitoring and observability; collaborate across teams, drive incident response and on-call rotation, capacity planning, and postmortems to improve reliability.
Top Skills: AnsibleCCi/CdDpdkGitopsGoHelmKubernetesLinuxMonitoring/ObservabilityOpenstack NeutronOvnOvsPythonSmartnicsSr-IovTerraform
YesterdaySaved
Remote
USA
122K-162K Annually
Mid level
122K-162K Annually
Mid level
Software
Provide senior-level technical escalation and support for GPU/HPC cloud infrastructure. Troubleshoot hardware, drivers, kernel, networking, and workload issues; perform root-cause analysis across clusters; build automations and documentation; mentor junior engineers; participate in on-call rotations and incident response; collaborate with engineering to implement permanent fixes.
Top Skills: AnsibleCi/CdCudaDatadogDockerFirewallsGpudirect RdmaGrafanaInfinibandKubernetesLinuxNcclNvidia GpusNvlinkPrometheusRoceSlurmTcp/IpTerraformVpn
Reposted 5 Days AgoSaved
In-Office or Remote
USA
125K-195K Annually
Senior level
125K-195K Annually
Senior level
Software
The Senior Incident Manager leads incident response for AI infrastructure, coordinates cross-team collaboration, oversees incident resolution and post-incident reviews to enhance operational resilience.
Top Skills: Cloud PlatformsDatadogGpu ClustersGrafanaJIRANetworkingPagerdutyPrometheusServicenow
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account