Ollama Inc. Logo

Ollama Inc.

Software Engineer, Runtime

Posted 2 Days Ago
In-Office
Palo Alto, CA, USA
Entry level
In-Office
Palo Alto, CA, USA
Entry level
Develop and optimize Ollama’s local runtime for open-model inference across macOS, Linux, and Windows. Responsibilities include model loading, scheduling, memory management, quantization, GPU backends, model architecture integrations, and performance improvements. The role involves Go and C/C++ development, close-to-the-metal systems work, profiling, hardware optimization, open-source collaboration, and partnerships with model labs and hardware vendors.
The summary above was generated by AI

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll work on the heart of Ollama — the local runtime that runs open models on developers' own machines. It loads models, manages memory, drives GPU acceleration across NVIDIA, AMD, Intel, Qualcomm, and Apple Silicon (including our MLX integration), and makes all of it feel instant. You'll work in Go and C/C++ and touch the model formats and inference engines underneath, shipping to macOS, Linux, and Windows across an enormous range of hardware.

What you'll do
  • Make open models run fast and reliably on consumer and enterprise hardware — from a MacBook Pro to server-grade NVIDIA GPUs.

  • Own pieces of the runtime: model loading & scheduling memory management, quantization, GPU hardware backends.

  • Integrate new model architectures and quantization formats so the latest open models work on day one.

  • Improve cold-start, time-to-first-token, and throughput

  • Partner with model labs and hardware vendors on early access and deep integrations.

  • Ship in the open: Ollama is open source, and you'll work with the community

Example projects
  • Add support for a new model family end-to-end — format parsing, weights loading, and the defaults that make it useful out of the box.

  • Cut cold-start for a popular model in half by streaming weights and lazy-loading layers.

  • Land a new quantization format so a 70B model runs on a single consumer GPU.

  • Wire up a new GPU backend and find a 2x throughput win with kernel selection and memory tuning.

  • Improve the "Auto" experience — picking the right model and settings for a machine's hardware without the user thinking about it.

You may be a fit if
  • You have strong systems fundamentals and are comfortable in Go, C, or C++

  • You've worked close to the metal — GPU compute, inference, game engines, databases, OS, or networking.

  • You care about performance and have profiled and optimized real workloads.

  • You're comfortable shipping to millions of users and handling the long tail of hardware and OS combinations.

  • Bonus: experience with model quantization, GPU programming (CUDA/Metal/SYCL), Apple MLX

Similar Jobs

18 Days Ago
Hybrid
Santa Clara, CA, USA
130K-220K Annually
Senior level
130K-220K Annually
Senior level
Software
Design and optimize production machine-learning runtime pipelines for autonomous vehicle motion planning. Build high-performance C++ inference, feature extraction, trajectory generation, and orchestration components for embedded platforms. Develop validation, safety guardrails, fallback mechanisms, testing, monitoring, debugging, and observability tools. Investigate complex runtime failures and optimize latency, throughput, memory, and multithreaded performance while collaborating with machine-learning and systems engineering teams.
Top Skills: C++C++17C++20Ci/CdCloud InfrastructureCudaEmbedded Gpu PlatformsOnnx RuntimePyTorchTensorrt
25 Days Ago
Hybrid
San Francisco, CA, USA
266K-445K Annually
Entry level
266K-445K Annually
Entry level
Artificial Intelligence • Machine Learning • Generative AI
Build low-level device runtime software for OpenAI custom AI accelerators. Responsibilities include kernel scheduling, command submission, queueing, device memory management, synchronization, hardware-software interfaces, simulation-based validation, performance tuning, debugging, testing, tracing, and profiling. Collaborate with compiler, kernel, architecture, verification, firmware, and silicon teams to improve accelerator execution and co-design.
Top Skills: CC++CachingCycle-Accurate SimulationDevice RuntimesDmaDriversEvent-Based SimulationFirmwareMemory CoherencyOperating SystemsRustVirtual Memory
One Month Ago
In-Office
San Francisco, CA, USA
144K-198K Annually
Mid level
144K-198K Annually
Mid level
Aerospace • Defense
Design, implement, and maintain onboard software packaging, deployment, and runtime infrastructure for satellites. Work across bundle management, Linux rootfs and service management, container runtime/environment and registries, CI/CD pipelines, and product validation to ensure reliable, secure execution of customer workloads in orbit.
Top Skills: ArmC++Cue (Cuelang)Device TreeDockerGitGitlab-CiGoGoogletestGrpcJenkinsKubernetesLinuxNatsPytestPythonRustSystemdUnityYoctoZmq

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account