Top Tech Jobs & Startup Jobs in San Francisco Bay Area, CA

Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and maintain agent APIs and production agent applications for document understanding, advanced RAG, and customer support automation. Integrate open-source models, collaborate with backend and infra for deployment and monitoring, and ensure APIs are robust, scalable, and developer-friendly.
Top Skills: Agent FrameworksAPIsCloud-Native DevelopmentContainer OrchestrationHugging FaceKubernetesLangchainLlamaindexMultimodal ModelsOcrPythonRagSummarization
Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Own and evolve core backend microservices for an AI inference platform: build production-grade APIs, multi-tenant SaaS features (auth, RBAC, billing), design OLTP/OLAP data models, collaborate on multi-cloud orchestration, ensure reliability/performance, and drive engineering quality through testing and CI/CD.
Top Skills: Ci/CdClickhouseFastapiGraphQLGrpcHugging FaceKubernetesLlm ServingNext.JsOpentelemetryPostgresPythonReactRestSQL
Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design and optimize high-performance GPU kernels (GEMM, attention, routing) for AI inference across NVIDIA and AMD GPUs. Implement CUDA/C++ and low-level assembly code, build reduced-precision/quantized (FP8/FP4) kernels, benchmark cross-vendor performance, contribute to internal GPU libraries, accelerate multi-modal pipelines, and integrate next-generation GPU features into production.
Top Skills: AmdC++CudaCutlassFp4Fp8GemmGpu AssemblyHipNvidiaRocmTriton
Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and maintain a scalable web platform and APIs for deploying and monitoring multimodal AI models and agent workflows. Collaborate with product, infrastructure, and design teams to optimize performance, ensure reliability, drive CI/CD and testing, and contribute to long-term architecture decisions for a cloud-native, multi-tenant SaaS system.
Top Skills: Ci/CdCloud-NativeFastapiGraphQLGrpcHugging FaceKubernetesLlmNext.JsOpentelemetryPostgresPythonRbacReactRestSQLTypescript
Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Lead end-to-end enterprise sales for FriendliAI's AI inference platform: generate pipeline, close high-value deals, run technical POCs, engage AI/ML communities, collaborate with engineering, and inform product roadmap.
Top Skills: Ai/MlCloud PlatformsDeveloper ToolsHugging FaceInference ServingLlmMlops
Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, deploy, and operate large-scale LLM and multimodal inference architectures. Work hands-on with customer engineering teams to containerize, scale, monitor, and troubleshoot GPU-based inference workloads across Kubernetes, CI/CD, and hybrid/on-prem environments. Create Helm charts, Terraform modules, and observability tooling while delivering workshops and platform reliability insights.
Top Skills: AWSCi/CdDeepspeed-InferenceDockerDocker ImagesEksElkGCPGpu ComputingGrafanaHelmHugging FaceKubernetesLokiOciOtelPrometheusTensorrtTerraformTritonVllm
Reposted 23 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, implement, and optimize GPU kernels, kernel compiler, memory planner, and runtime for low-latency generative AI inference. Analyze performance bottlenecks across hardware and software, collaborate with infrastructure teams, and maintain production profiling, benchmarking, and validation tooling while supporting new model architectures and multi-GPU strategies.
Top Skills: BenchmarkingC++Compiler InfrastructureDiffusion ModelsDistributed InferenceGpu KernelsKernel CompilerMulti-GpuProfilingPythonRuntime SystemsTransformer Models
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account