d-Matrix Logo

d-Matrix

Senior Staff LLM Inference Engineer

Reposted 16 Days Ago
Be an Early Applicant
Hybrid
Santa Clara, CA, USA
195K-285K Annually
Expert/Leader
Hybrid
Santa Clara, CA, USA
195K-285K Annually
Expert/Leader
Lead end-to-end LLM inference engineering: prototype and deploy optimized inference systems across heterogeneous hardware. Build POCs, implement custom kernels and operator optimizations, drive quantization, sparsity, and batching strategies, maintain runtimes and serving frameworks, and contribute to distributed inference and performance profiling while collaborating with hardware, product, and business teams.
The summary above was generated by AI

At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution.  Ready to come find your playground? Together, we can help shape the endless possibilities of AI. 

D-Matrix Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack: from pathfinding emerging use cases and novel deployment patterns to deep optimization of inference kernels, to building proof-of-concept systems that showcase D-Matrix’s unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly with product and hardware teams to shape the roadmap.

We build the tools, runtimes, and frameworks that let frontier AI models run efficiently and cost-effectively across heterogeneous deployments — combining D-Matrix silicon with CPUs, GPUs, and custom accelerators. Our work powers everything from benchmarking and evaluation pipelines to production-grade inference serving.

This Role

We are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack — from kernel-level optimization to distributed orchestration to high-level serving APIs.

This role could be a great match for you if you:

• Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time.

• Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed.

• Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas.

• Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff.

• Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used.

• Value clear communication and thrive in a small, high-ownership team environment.

Responsibilities

• Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.

• Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.

• Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.

• Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.

• Build and maintain inference runtimes, serving frameworks, and evaluation tooling.

• Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.

• Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.

• Partner with product and business development to translate POCs into customer-facing demonstrations.

• Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.

Required Qualifications

• Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience.

• Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.

• Strong proficiency in Python and C/C++.

• Hands-on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).

• Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.

• Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.

Preferred Qualifications

• Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).

• Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments).

• Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.

• Contributions to open-source inference or ML systems projects.

• Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).

• Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.

• Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.

Why D-Matrix Frontier Group

• Work on genuinely novel hardware — D-Matrix in-memory compute architecture opens up inference optimization problems that don’t exist anywhere else.

• End-to-end ownership from idea to deployed system, with a short feedback loop between your work and real hardware.

• Small, senior team with high autonomy and direct influence on product direction.

• Competitive compensation, equity, and benefits in Santa Clara, CA.

 

Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation.

HQ

d-Matrix Santa Clara, California, USA Office

5201 Great America Pkwy, Santa Clara, CA, United States, 95054

Similar Jobs

36 Minutes Ago
In-Office
San Francisco, CA, USA
96K-140K Annually
Senior level
96K-140K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead month-end close for assigned areas, manage GL cycles (AR/revenue, accruals, fixed assets, deferred revenue, equity), ensure reconciliations and GL integrity, drive controls and audit readiness, coordinate cross-functional cutoffs and analytics, and improve accounting processes and automation.
Top Skills: ExcelFloqastGoogle Sheets
37 Minutes Ago
In-Office
88K-120K Annually
Senior level
88K-120K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Own month-end close for assigned areas, prepare journals and reconciliations, manage AR/revenue, accruals, fixed assets, deferred revenue and equity cycles, improve controls and audit readiness, streamline processes and reconciliation tooling, and coordinate analytics with cross-functional teams.
Top Skills: ExcelFloqastGl SystemsGoogle SheetsReconciliation ToolsSubledger Systems
51 Minutes Ago
In-Office
119K-161K Annually
Senior level
119K-161K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Lead integration and implementation of specialized test equipment (STE) for satellite production. Develop manufacturing/test processes, resolve complex production issues, perform data/trend analysis, coordinate staffing and STE allocation, drive corrective/preventive actions, interface with internal/external stakeholders, and support long-range lab capability planning and sustainment.
Top Skills: ElectronicsMechanical SystemsSatellite Test EquipmentSpecialized Test Equipment (Ste)

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account