MakerMaker.AI Logo

MakerMaker.AI

KERNEL ENGINEER

Posted 3 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Mid level
In-Office
San Francisco, CA, USA
Mid level
Write and optimize GPU kernels for training and inference, profile workloads with hardware counters, co-design kernels with researchers, integrate and benchmark kernels in training/serving stacks, and maintain kernel quality while sharing expertise across the team.
The summary above was generated by AI

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You'll write and optimize the GPU kernels and supporting systems software that makes our training and inference workloads fast. This is deep, low-level work (performance counters, memory bandwidth, warp-level scheduling) applied to the specific shapes and patterns our models actually use.

We hire kernel engineers because the gap between "this works" and "this is fast on the hardware we have" is enormous, and that gap directly bounds what our researchers can try. You'll close that gap.

WHAT YOU'LL DO

  • Write and optimize GPU kernels (CUDA, ROCm, Triton, or similar) for training and inference workloads: attention variants, MoE layers, custom activations, communication primitives

  • Profile real workloads with hardware counters and translate findings into specific kernel-level optimizations

  • Co-design kernels with the research teams, when the kernel and the algorithm need to change together, you participate in both

  • Integrate optimized kernels into our training and serving stacks; benchmark before and after; verify the win is real end-to-end

  • Maintain kernel quality over time as hardware, frameworks, and workloads shift underneath

  • Spread kernel-level fluency across the team; we want this expertise shared, not siloed

WHAT WE'RE LOOKING FOR

  • 4+ years writing performant GPU kernels (CUDA, ROCm, Triton, or production-grade equivalent)

  • Hardware-level fluency: memory hierarchy, occupancy, register pressure, tensor cores, warp scheduling

  • Profiling fluency (Nsight, ncu, or comparable tools) and the discipline to measure before changing

  • Track record of shipping kernel-level optimizations that moved a measurable metric in a real system

  • Strong systems expertise: you understand how kernels live inside larger frameworks and how integration choices affect end-to-end performance

  • Comfortable reading framework-level Python and C++ around your kernels

NICE TO HAVE

  • Open-source contributions to kernel libraries, compilers, or ML frameworks

  • Experience with multiple accelerator architectures (different GPU families, TPUs, custom ASICs), preferably AMD GPUs

  • Familiarity with collective communication primitives (NCCL or equivalent)

  • Compiler or runtime background

THIS ROLE IS PROBABLY NOT FOR YOU IF

  • You haven't gotten your hands dirty at the kernel level: this isn't a higher-level systems role rebranded

  • You want to stay narrowly in one library; we expect breadth across the kernel surface our models actually use

  • Performance work without measurable end-to-end impact frustrates you

Similar Jobs

17 Days Ago
In-Office
Sunnyvale, CA, USA
182K-242K Annually
Senior level
182K-242K Annually
Senior level
Cloud • Information Technology • Machine Learning
Author, profile, and optimize CUDA GPU kernels for LLM inference to maximize throughput and minimize latency. Build reproducible microbenchmarks and MLPerf workflows, collaborate across product, hardware, and orchestration teams, lead designs and reviews, mentor engineers, and translate kernel wins into end-to-end model-serving performance improvements.
Top Skills: C++CudaLlm-DMlperfNsight ComputeNsight SystemsNvlinkPciePythonSglangTensor CoresTensorrt-LlmVllm
9 Hours Ago
In-Office
Walnut Creek, CA, USA
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Automotive
Develop and maintain core OS kernel components for the VxWorks RTOS. Design, implement, test, and debug kernel features and device drivers; create design documents and code reviews; collaborate with cross-functional teams for hardware/software integration; follow CI/CD and agile practices; investigate verification and customer issues; drive projects end-to-end ensuring performance, reliability, and security.
Top Skills: AgileAspiceAssemblyBspCC++Ci/CdDevice DriversDo-178CEmbedded SystemsIso 26262MilPythonRtosVxworks
4 Days Ago
In-Office
Santa Clara, CA, USA
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop silicon-accurate GPU kernel microbenchmarks, model-level performance analysis, and agentic kernel optimization to maximize LLM inference throughput and latency. Attribute bottlenecks across kernel, compiler, and runtime, produce optimization policies, and collaborate with compiler, hardware, kernel, and framework teams to deliver production-grade performance improvements.
Top Skills: C++CudaCuptiCutlassNcuNsysPtxPythonSassSglangTritonTrt-LlmVllm

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account