Luma AI Logo

Luma AI

Senior Machine Learning Engineer - Hardware Abstractions & Performance Optimization

Sorry, this job was removed at 01:04 p.m. (PST) on Thursday, Jul 24, 2025
Remote
Hiring Remotely in United States
220K-300K Annually
Remote
Hiring Remotely in United States
220K-300K Annually

Similar Jobs

21 Minutes Ago
Remote
United States
85K-95K Hourly
Entry level
85K-95K Hourly
Entry level
Software
Lead go-to-market for a new AI product by conducting customer discovery, defining positioning and journeys, improving activation and adoption, building scalable GTM systems, shaping product roadmap with engineering, and running experiments to optimize conversion, engagement, and retention.
49 Minutes Ago
Easy Apply
Remote or Hybrid
3 Locations
Easy Apply
145K-208K Annually
Senior level
145K-208K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Manage major healthcare enterprise accounts in UT/AZ/CA; build C-suite relationships; develop long-term account strategies; coordinate internal teams; act as trusted advisor aligning Zscaler cloud security solutions to client goals; drive new logo acquisition and revenue targets.
Top Skills: AICloud-NativeZero Trust ExchangeZscaler
49 Minutes Ago
Remote
USA
175K-225K Annually
Senior level
175K-225K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead strategy, roadmap, and business performance for a core AI infrastructure product area. Drive customer research, KPI-driven roadmaps, business cases, cross-functional launches, and measurable outcomes for adoption, retention, revenue, and platform expansion.
Top Skills: Ai/MlAPIsCloud PlatformsDeveloper PlatformsDistributed SystemsUsage-Based Compute

Luma’s mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.

We are looking for engineers with significant experience maintaining & designing highly efficient systems and code that can be optimized to run on multiple hardware platforms, bringing our state-of-the-art models to as many people at the best performance per dollar.

Responsibilities
  • Ensure efficient implementation of models & systems with a focus on designing, maintaining, and writing abstractions that scale beyond NVIDIA/CUDA hardware.

  • Identify and remedy efficiency bottlenecks (memory, speed, utilization, communication) by profiling and implementing high-performance PyTorch code, deferring to Triton or similar kernel-level languages as necessary.

  • Benchmarking our products across a variety of hardware & software to help the product team understand the optimal tradeoffs between latency, throughput and cost at various degrees of parallelism.

  • Work together with our partners to help them identify bottlenecks and push forward new iterations of hardware and software.

  • Work closely together with the rest of the research team to ensure systems are planned to be as efficient as possible from start to finish and raise potential issues for hardware integration.

Must have experience
  • Experience optimizing for memory, latency and throughput in Pytorch.

    • Bonus: experience with non-NVIDIA systems

  • Experience using torch.compile / torch.XLA.

  • Experience benchmarking and profiling GPU & CPU code in Pytorch for optimal device utilization (examples: torch profiler, memory profilers, trace viewers, custom tooling).

  • Experience building tools & abstractions to ensure models run optimally on different hardware and software stacks .

  • Experience working with transformer models and attention implementations.

  • Experience with parallel inference, particularly with tensor parallelism, pipeline parallelism.

Good to have experience
  • Experience with high-performance Triton/CUDA and writing custom PyTorch kernels and ops. Top candidates will be able to write fused kernels for common hot paths, understand when to make use of lower level features like tensor cores or warp intrinsics, and will understand where these tools can be most impactful.

  • Experience writing high-performance parallel C++. Bonus if done within an ML context with PyTorch, like for data loading, data processing, inference code

  • Experience building inference / demo prototype code (incl. Gradio, Docker etc.)

HQ

Luma AI San Francisco, California, USA Office

San Francisco, CA, United States

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account