Top Tech Jobs & Startup Jobs in San Francisco Bay Area, CA

6 Hours AgoSaved
Remote or Hybrid
San Francisco, CA, USA
160K-220K Annually
Senior level
160K-220K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Leads complex, cross-functional infrastructure programs spanning compute, networking, storage, observability, hardware, and datacenter operations. Owns program planning, milestones, dependencies, risks, and execution across multiple workstreams. Partners with engineering, operations, security, platform, and leadership teams to deliver network expansions, datacenter growth, hardware rollouts, GPU platform deployments, and infrastructure readiness. Provides leadership updates and improves execution frameworks, tooling, and program visibility.
Top Skills: Cloud InfrastructureCompute InfrastructureDatacenter InfrastructureDistributed SystemsGpu ArchitecturesHardware PlatformsNetwork InfrastructureObservabilityPytorch LightningStorage Systems
6 Hours AgoSaved
Remote or Hybrid
San Francisco, CA, USA
165K-310K Annually
Entry level
165K-310K Annually
Entry level
Artificial Intelligence • Machine Learning • Software
Develop and post-train deep learning models while building systems, tooling, and workflows for training, evaluation, debugging, and deployment. Contribute to open-source projects, distributed AI infrastructure, backend services, and developer platforms. Collaborate with research, product, infrastructure, and customer teams to solve complex technical problems, prototype ideas, and productionize successful experiments.
Top Skills: Cloud InfrastructureCudaDeepspeedDistributed SystemsFsdpHugging FaceLightning AiNvidia MoltPyTorchSglangTritonVllm
6 Hours AgoSaved
Hybrid
San Francisco, CA, USA
160K-275K Annually
Senior level
160K-275K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Own the vision, roadmap, launch, adoption, pricing, and go-to-market for Lightning AI’s experimentation and post-training products. Define workflows for experiments, fine-tuning, reinforcement learning, evaluation, model comparison, and production handoffs. Collaborate closely with engineering, customers, executives, sales, Growth, and Finance. Drive product strategy, technical requirements, model evaluations, pricing, metrics, developer experience, and end-to-end execution for ML infrastructure and developer tooling.
Top Skills: APIsDistributed TrainingExperiment TrackingHugging FaceHyperparameter OptimizationKubernetesMlflowModel EvaluationModel RegistriesObservabilityPreference OptimizationPyTorchPytorch LightningRayReinforcement LearningSdksSlurmSupervised Fine-TuningWeights & Biases
6 Hours AgoSaved
Hybrid
San Francisco, CA, USA
180K-250K Annually
Senior level
180K-250K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Design and scale backend systems for AI agent orchestration, distributed workflows, experiment management, platform APIs, and developer tooling. Build reliable cloud-native infrastructure, improve observability and performance, partner across engineering and research teams, maintain software quality through testing and continuous delivery, and mentor engineers on distributed systems and backend best practices.
Top Skills: APIsAsynchronous ProcessingAutomated TestingAWSAzureCloud-Native ArchitectureContinuous DeliveryDistributed SystemsDockerGCPGoKubernetesObservabilityPythonRust
6 Hours AgoSaved
Hybrid
San Francisco, CA, USA
120K-250K Annually
Entry level
120K-250K Annually
Entry level
Artificial Intelligence • Machine Learning • Software
Develop and scale the Lightning AI platform across frontend, CLI, APIs, and backend systems. Build features using React, Python, or Go; improve stability, performance, architecture, automation, and continuous delivery. Collaborate with engineering, product, and design leaders, own end-to-end feature development, reduce technical debt, and mentor engineers on system design and problem-solving.
Top Skills: Go (Golang)PythonPytorch LightningReactSaaS
6 Hours AgoSaved
Hybrid
San Francisco, CA, USA
115K-140K Annually
Entry level
115K-140K Annually
Entry level
Artificial Intelligence • Machine Learning • Software
Support ML engineers running large-scale training and inference workloads across Kubernetes, cloud, and GPU infrastructure. Diagnose distributed PyTorch, CUDA, networking, storage, scheduling, performance, and reliability issues. Analyze observability data, guide customers through complex infrastructure problems, support production incidents, and improve platform reliability through automation, tooling, documentation, runbooks, and operational processes. The role partners closely with infrastructure, networking, and platform engineering teams in a hybrid office environment.
Top Skills: Bare Metal InfrastructureCloud InfrastructureCudaDistributed SystemsDockerGpu InfrastructureGrafanaInference ServingInfinibandKubeflowKubernetesLinuxMl InfrastructureNcclOpentelemetryPrometheusPythonPyTorchRayRdmaSlurmStorage Systems
6 Hours AgoSaved
Hybrid
San Francisco, CA, USA
120K-250K Annually
Entry level
120K-250K Annually
Entry level
Artificial Intelligence • Machine Learning • Software
Develop and scale the Lightning AI platform’s frontend and UI infrastructure using React and Redux. Build end-to-end features, improve stability and performance, evaluate technical architecture, automate software delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a rapidly changing SaaS environment while maintaining high standards for code quality and system design.
Top Skills: JavaScriptReactReduxTypescript
6 Hours AgoSaved
Remote or Hybrid
USA
150K-190K Annually
Senior level
150K-190K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Design and scale reliable spine-leaf data center networks supporting GPU clusters and AI workloads. Build Ethernet fabrics using EVPN/VXLAN and BGP, optimize traffic, support RoCE/RDMA, and manage backbone, WAN, DCI, and edge connectivity. Develop network automation and Infrastructure-as-Code solutions, improve observability and telemetry, troubleshoot distributed performance issues, and collaborate with compute, storage, platform, and operations teams. The role requires hands-on Cumulus NOS expertise and experience with large-scale, high-performance infrastructure.
Top Skills: AnsibleBgpCloud ConnectCumulus LinuxCumulus NosDirect ConnectEthernetEvpnInfinibandInfrastructure As CodeJunosNfvNvidia BluefieldNvidia QuantumNvidia SpectrumPythonRdmaRoceSonicTerraformVpcVxlan
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account