General Compute Logo

General Compute

Head of Infrastructure

Posted 2 Days Ago
Be an Early Applicant
Hybrid
San Francisco, CA, USA
Senior level
Hybrid
San Francisco, CA, USA
Senior level
Own the infrastructure layer of an AI inference cloud, including the Kubernetes-based control plane, gateway, load balancing, observability, capacity planning, accelerator operations, and incident response. The role begins hands-on and expands to building and managing the infrastructure team. Responsibilities include supporting heterogeneous ASIC and GPU fleets, partnering with hardware vendors, optimizing tail latency, establishing on-call practices, and bringing up new hardware platforms.
The summary above was generated by AI
About Us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.

About the role

You'll own the infrastructure layer of our inference cloud end-to-end. Today that means the control plane, the gateway in front of our ASIC fleet, and the observability stack that tells us where every millisecond goes. Over the next 6-8 months, it will grow into a heterogeneous fleet: ASICs for decode, GPUs for pre-fill, and the physical-layer ownership that comes with it.

The first six months are hands-on: k8s manifests, dashboards, oncall, and a direct line to our ASIC partner's engineering team when production behaves strangely. The team grows under you from there.

What you'll do:
  • Own the inference control plane. At the moment, it's built on configuration provided by our ASIC partner; you'll be the person who understands it deeply enough to modify, extend, and eventually replace pieces of it.

  • Own the gateway and load balancer that fronts the fleet. Model placement, request routing, and tail-latency engineering live here, driven by live utilization and per-model SLOs.

  • Own observability end-to-end. Per-request tracing from OpenRouter ingress through to the accelerator, with p50/p95/p99 dashboards, SLOs, and alerting that wakes the right person.

  • Run capacity planning against a real, distributed traffic mix across the open-weight models we serve.

  • Own the operational side of the ASIC partnership. Most weird production issues route through their engineering team until we build that expertise in-house, and you'll be our technical face in those conversations.

  • Bring up the pre-fill side of our disaggregated architecture on a second hardware platform as it comes online. Different vendor, different fabric, different kernels.

  • Build the on-call and incident response practice from zero. Hire and grow the team underneath you.

What we need from you:
  • 7+ years in infrastructure, SRE, or platform engineering, with at least some of it at a serious inference, ML, or HPC shop.

  • Hands-on with Kubernetes at production scale — not just deploying, but debugging the weird stuff.

  • Strong instincts for tail latency. You think about p99 and utilization as the same problem, not different ones.

  • Comfortable owning a vendor relationship where the vendor's bugs are now your production issues.

  • Track record of building observability practices that actually catch problems, not just generate dashboards.

  • Have been on-call through real incidents and can talk about what you learned.

  • Want to be the first infra hire at something early, not the tenth at something big.

NIce-to-Haves:
  • Experience operating non-NVIDIA accelerators in production — TPUs, ASICs, or alternative GPU vendors.

  • Background with model-serving stacks (vLLM, TGI, TensorRT-LLM, SGLang).

  • Network fabric experience at data-center scale (RoCE, InfiniBand).

  • Have hired and managed an infra team before.

  • Comfort at the hardware boundary — firmware, drivers, thermals — for when the roadmap takes us there.

Similar Jobs

3 Days Ago
In-Office or Remote
2 Locations
272K-489K Annually
Expert/Leader
272K-489K Annually
Expert/Leader
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Leads NVIDIA’s Infrastructure Security Engineering organization for EDA clusters. Builds and operates security controls across compute, storage, and networking in on-premises and multi-cloud environments while preserving performance for long-running workloads. Owns the technical roadmap, security technology lifecycle, automation, instrumentation, compliance evidence, and cross-functional execution. Develops secure-by-default infrastructure and leads teams responsible for scalable enforcement, observability, mitigation, and security outcomes.
Top Skills: CnappCspmDpuEbpfEdr/XdrIamInfrastructure-As-CodeIso 27001Multi-CloudPolicy-As-CodeSmartnicSoc 2
3 Days Ago
In-Office or Remote
2 Locations
272K-489K Annually
Expert/Leader
272K-489K Annually
Expert/Leader
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Leads a globally distributed customer engineering organization supporting internal EDA infrastructure while overseeing a Kubernetes platform team. Responsibilities include developing managers, improving support operations, defining SLOs, translating support trends into engineering roadmaps, automating recurring operational work, and advancing AI/LLM-based support assistants. The role requires hands-on technical engagement with Kubernetes, EDA job schedulers, internal LLM infrastructure, and production environments.
Top Skills: Ai/MlAutomation PipelinesEda InfrastructureKubernetesLlm InfrastructureLsfSlosSlurm
One Month Ago
In-Office
San Leandro, CA, USA
160K-250K Annually
Expert/Leader
160K-250K Annually
Expert/Leader
Energy • Renewable Energy
Lead Fuse's capital development: manage permitting, contractors, schedules, budgets, and regulatory navigation for current facility and greenfield builds. Coordinate architects, structural/MEP engineers, and agencies; hire and scale a capital projects team; remove permitting bottlenecks and deliver fast, compliant project execution while reporting risks and progress to the CEO.
Top Skills: Building CodesDoeFire CodesHazmatItarMepNfpa 30NrcPulsed PowerRadiation Facility RequirementsZoning

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account