Top Tech Jobs & Startup Jobs in San Francisco Bay Area, CA

7 Days AgoSaved
In-Office
2 Locations
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Designs and evolves large-scale microservices ecosystems using cloud-native platforms and distributed systems patterns. Leads service decomposition, domain-driven design, distributed data management, resilience, observability, security, API and event-driven architecture. Operates high-traffic production systems, guides platform engineering and cloud cost optimization, and leads architecture initiatives across multiple teams while managing stakeholders.
Top Skills: C#Cloud PlatformsContainer OrchestrationEnvoyGoIstioJavaKubernetesLinkerdMicroservicesService Mesh
7 Days AgoSaved
In-Office
Foster City, CA, USA
100K-110K Annually
Senior level
100K-110K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Designs, builds, and operates secure, highly available AWS cloud platforms and production workloads. Responsibilities include AWS architecture, infrastructure as code, automation, Kubernetes or ECS operations, CI/CD, cloud security, observability, troubleshooting, cost optimization, and continuous operational improvement. The engineer collaborates with application, security, and SRE teams to deliver resilient, scalable cloud-native solutions.
Top Skills: Amazon EcsAmazon EksAWSAws CdkAws OrganizationsBashCi/CdCloudFormationCloudfrontEbpfEc2FedrampFinopsGoHipaaIamLambdaPci-DssPowershellPythonRdsS3Service MeshSoc 2TerraformVpcZero-Trust Networking
7 Days AgoSaved
In-Office
Foster City, CA, USA
100K-100K Annually
Senior level
100K-100K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Optimize large-scale AI training and inference systems for throughput, latency, and cost. Responsibilities include GPU kernel and compiler optimization, distributed system tuning, memory and parallelism optimization, profiling across CPU and GPU systems, model compression, and production-scale LLM inference improvements. The role requires rigorous measurement, debugging, cross-functional collaboration, code and design reviews, mentorship, and delivery of reliable production engineering solutions.
Top Skills: C++Cpu ProfilingCutlassDeep LearningDeepspeedDistributed SystemsFinopsGpu ProfilingGpusHigh-Performance ComputingMachine LearningModel CompressionModel ParallelismPythonTensorrt-LlmTritonVllm
7 Days AgoSaved
In-Office
Foster City, CA, USA
105K-143K Annually
Senior level
105K-143K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Design, build, and operate reliable, high-performance inference platforms for large machine learning models in production. Responsibilities include request routing, batching, caching, autoscaling, GPU utilization, observability, distributed systems, performance engineering, capacity planning, incident response, and cost-efficient serving across diverse workloads.
Top Skills: AutoscalingC++Cloud PlatformsGoGpuKubernetesMetricsModel CompressionModel DistillationModel QuantizationPythonRustStructured LoggingTensorrt-LlmTracingVllm
7 Days AgoSaved
In-Office
2 Locations
70K-110K Annually
Senior level
70K-110K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Provides technical leadership across teams, drives complex engineering initiatives, shapes architecture, and executes technical roadmaps. Responsibilities include translating ambiguous requirements into scalable solutions, conducting code and design reviews, mentoring engineers, improving reliability and observability, and influencing engineering standards and culture. The role requires deep expertise in distributed systems, cloud platforms, or large-scale data systems, along with strong backend programming, system design, communication, and cross-functional collaboration skills.
Top Skills: Backend ProgrammingCloud PlatformsDistributed SystemsFinopsLarge-Scale Data SystemsObservabilitySystem Architecture
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
7 Days AgoSaved
In-Office
Foster City, CA, USA
85K-110K Annually
Senior level
85K-110K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Develop and optimize CUDA-based workloads for AI, scientific computing, inference, and high-throughput processing. Design custom GPU kernels, profile production workloads, integrate kernels into ML frameworks, and improve performance across modern accelerator platforms. Collaborate with engineering and research teams, participate in code and design reviews, translate requirements into production solutions, and mentor junior engineers.
Top Skills: Cuda C/C++CutlassDistributed Training InfrastructureFastertransformerGpu ArchitectureGpu ProgrammingHigh-Performance ComputingHigh-Performance InterconnectsLlvmMachine Learning FrameworksMlirMpiNcclTensorrtTritonVllm
7 Days AgoSaved
In-Office
Foster City, CA, USA
90K-115K Annually
Senior level
90K-115K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Designs and operationalizes fine-tuning workflows for large language models using supervised, preference-based, and reinforcement learning. Responsibilities include dataset construction, evaluation methodology, distributed training, GPU cluster operations, failure recovery, and production pipeline reliability. The engineer collaborates with cross-functional teams, ships impactful LLM solutions, conducts design and code reviews, and mentors junior engineers.
Top Skills: DpoFsdpGpu ClustersPipeline ParallelismPythonPyTorchRlhfTransformer-Based Language ModelsZero
7 Days AgoSaved
In-Office
2 Locations
100K-120K Annually
Senior level
100K-120K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Designs, deploys, and operates secure, scalable Microsoft Azure infrastructure for production workloads. Responsibilities include cloud architecture, infrastructure as code, automation, AKS operations, CI/CD, security hardening, cost optimization, monitoring, observability, troubleshooting, and compliance. The engineer partners with application development, security, and SRE teams to deliver resilient cloud-native platforms.
Top Skills: Arm TemplatesAzure DevopsAzure Kubernetes Service (Aks)BashBicepGithub ActionsIstioLinkerdAzurePowershellPythonTerraform
7 Days AgoSaved
In-Office
2 Locations
75K-100K Annually
Senior level
75K-100K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Designs and builds end-to-end Industrial IoT systems connecting factory-floor devices to enterprise cloud platforms. Responsibilities include edge data acquisition, secure connectivity, cloud ingestion, real-time analytics, operational dashboards, OT/IT integration, troubleshooting, code and design reviews, production delivery, and mentoring junior engineers. The role requires programming, industrial protocol, cloud IoT, edge runtime, streaming analytics, and industrial cybersecurity expertise.
Top Skills: AWSAzureC++Digital-Twin PlatformsGCPGoGreengrassIec 62443InfluxdbIot EdgeIsaModbusMqttOpc UaPredictive Maintenance Machine LearningPythonReal-Time AnalyticsSparkplug BStreaming AnalyticsTimescaledbUnified Namespace
7 Days AgoSaved
In-Office
Foster City, CA, USA
80K-100K Annually
Senior level
80K-100K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Design, train, evaluate, and deploy reinforcement learning systems for complex decision-making problems. Develop reward functions, simulation environments, and neural network policies; scale training on GPU clusters; and move RL solutions from research into reliable production systems. The role requires strong theoretical and engineering expertise, with preferred experience in RLHF, multi-agent or hierarchical RL, robotics, autonomous driving, and open-source or published RL work.
Top Skills: Deep Learning FrameworksGpu ClustersLarge Language ModelsPythonReinforcement Learning LibrariesRlhfSimulation Environments
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account