Top Tech Jobs & Startup Jobs in San Francisco Bay Area, CA

6 Days AgoSaved
Remote
United States
Entry level
Entry level
Artificial Intelligence • Hardware • Software • Semiconductor
Develop kernel-centric reliability solutions for AI compute clusters and production services. Responsibilities include debugging failures, building diagnostic tools, supporting incident response, performing root-cause analysis, improving kernel and software reliability, and collaborating with systems, hardware, ASIC, and architecture teams on reliability-focused designs.
Top Skills: CC++Core Dump HandlingDebuggersDistributed ProgrammingGpusParallel ProgrammingProfilersPythonSanitizersTracing
One Month AgoSaved
Remote
United States
Expert/Leader
Expert/Leader
Artificial Intelligence • Hardware • Software • Semiconductor
Lead architecture and evolution of enterprise, data center, and cloud networks for hyperscale AI; design secure, high-performance fabrics (spine/leaf, RDMA); build AI-agent review frameworks; implement segmented zero-trust network designs; set resiliency, observability, and capacity standards; mentor global engineers and guide cross-functional initiatives.
Top Skills: Ai AgentsAnsibleAws Direct ConnectAws Transit GatewayAws VpcBgpEvpnHigh-Bandwidth FabricsInfinibandKubernetes (K8S)Kubernetes Service MeshRdmaRoceSpine-LeafTerraformVxlanZero Trust
One Month AgoSaved
Remote
US
Senior level
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate firewall, segmentation, and zero-trust controls across data center, corporate, and cloud networks. Implement infrastructure-as-code for policy and CI/CD deployment, manage lifecycle and rule hygiene, build network detection capabilities, operate VPN/remote access patterns, and document architecture and runbooks while partnering with Network Engineering and Security Operations.
Top Skills: AnsibleAWSCi/CdCloud-Native FirewallsDnsPythonRoutingSecurity GroupsSwitchingTcp/IpTerraformTlsTransit GatewayVpcVpnZtna
One Month AgoSaved
Remote
United States
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Join IT & Security to secure and scale enterprise IT, cloud, network, and infrastructure; build automation and tooling; support detection, response, and vulnerability management; improve identity, endpoint, and systems security; and partner cross-functionally to develop processes that enable secure, reliable, and scalable AI workloads.
Top Skills: AutomationCloud SecurityEndpoint SecurityIdentity And Access ManagementIncident ResponseInfrastructure-As-CodeNetwork SecurityScriptingSecurity OperationsVulnerability Management
One Month AgoSaved
Remote
United States
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills: Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
One Month AgoSaved
Remote
United States
Senior level
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement system-level debugging, validation, and observability platforms. Build automated anomaly collection/analysis, visualization and root-cause tools, failure classification and monitoring frameworks. Extend compilers, runtimes and programming interfaces for profiling and instrumentation, improve bring-up and low-level debug workflows, lead cross-functional initiatives, support incident response, and establish debuggability and reliability best practices.
Top Skills: C++CompilersFirmwareHardware InterfacesInstrumentationProfilingProgramming InterfacesPythonRuntimesVisualization Tools
One Month AgoSaved
Remote
USA
Senior level
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Develop, optimize, and validate high-performance ML and linear-algebra kernels for Cerebras wafer-scale hardware. Implement low-level assembly and a C-like DSL (CSL), model and analyze performance, build testing methodologies, and collaborate with chip and system architects to maximize compute utilization for state-of-the-art AI and HPC workloads.
Top Skills: AssemblyC++Cerebras WseCsl (C-Like Domain Specific Language)Distributed Memory SystemsFpgasGpusHpcMachine Learning FrameworksParallel AlgorithmsPythonPyTorchTensorFlow
One Month AgoSaved
Remote
United States
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, optimize, and validate high-performance ML and linear algebra kernels for Cerebras hardware. Develop low-level assembly and CSL routines, use mathematical performance models, create unit/system tests, and collaborate with chip and system architects to maximize compute utilization and scale kernels for state-of-the-art AI/HPC workloads.
Top Skills: AssemblyC++Cerebras WseCslDistributed Memory SystemsFpgasGpusHpcLinear AlgebraParallel AlgorithmsPythonPyTorchTensorFlow
One Month AgoSaved
Remote
United States
Senior level
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement high-performance distributed runtime components for large-scale training and inference. Optimize data and communication pipelines, enable scalable multi-node execution, collaborate with ML and compiler teams, diagnose performance issues via profiling, and contribute to system architecture and roadmap for cutting-edge AI workloads.
Top Skills: CC++Cluster ComputingCompilersDistributed SystemsHpcInter-Process CommunicationMemory ManagementMulti-ThreadingNetworkingProfiling/InstrumentationPythonPyTorch
One Month AgoSaved
Remote
United States
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Integrate, validate, and productionize cross-stack inference features across AI frameworks, runtime, compiler, kernels, distributed systems, and hardware. Drive zero-to-one projects, debug system-wide failures, manage accelerated timelines, and improve automation, diagnostics, and repeatable integration practices while collaborating across software and hardware teams.
Top Skills: Ai FrameworksC++Cloud InfrastructureCluster OrchestrationCompilersContainersDistributed SystemsGoHigh-Performance ComputingKernelsLlmsMicroservicesObservabilityPerformance DebuggingProfilingPythonRuntimes
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account