NVIDIA Jobs

Deep Learning Architect, LLM Inference - New College Grad 2026

Sorry, this job was removed at 08:27 p.m. (PST) on Tuesday, Jun 02, 2026

Be an Early Applicant

In-Office

Santa Clara, CA, USA

In-Office

Santa Clara, CA, USA

Similar Jobs

Adyen

Compliance Advisory Officer

41 Minutes Ago

Easy Apply

Hybrid

San Francisco, CA, USA

Easy Apply

120K-155K Annually

Senior level

120K-155K Annually

Senior level

Fintech • Payments • Financial Services

Support assessment and resolution of escalated compliance matters, analyze AML/CFT and integrity risks, partner with commercial and first-line teams to apply risk-based solutions, help develop compliance frameworks and escalation procedures, identify automation opportunities, and collaborate globally to ensure compliant onboarding and operations.

Adyen

Compliance Advisory Officer

41 Minutes Ago

Easy Apply

Hybrid

San Francisco, CA, USA

Easy Apply

120K-155K Annually

Senior level

120K-155K Annually

Senior level

Fintech • Payments • Financial Services

Support assessment and resolution of escalated compliance matters across AML/CFT, integrity, and regulatory obligations. Partner with legal, risk, commercial, and operations to provide risk-based solutions, develop compliance frameworks, improve escalation procedures, and identify automation opportunities. Translate compliance issues into actionable steps and collaborate globally to execute compliance initiatives.

Tapestry - Coach and Kate Spade

Retail Contingent

2 Hours Ago

Hybrid

Livermore, CA, USA

15-24 Hourly

Entry level

15-24 Hourly

Entry level

eCommerce • Fashion • Retail • Sales • Wearables • Design

Maintain organized, customer-ready store by processing deliveries, stocking the sales floor, executing price changes and markdowns, auditing inventory/shrinkage, and supporting daily operational standards and cleanliness.

Top Skills: Omnichannel SellingSocial Media

We are now looking for a Deep Learning Architect, LLM Inference!

NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS, then this is a great role for you!

What you'll be doing:

You will do workload characterization of the latest LLMs and inference servers like vLLM, SGLang and TRT-LLM to ensure NVIDIA maintains its leadership position.
Join forces with the performance marketing team to build engaging content, including blog posts and updates to InferenceX to highlight NVIDIA's outstanding inference achievements.
Collaborate with engineers from AI startup companies to establish standard benchmarking methodologies.
Develop a constantly evolving inference performance data results website.
Invent E2E profiling and analysis tools that you will use to keep up with the rapid pace of Generative AI.
Contribute to deep learning software projects, such as PyTorch, TRT-LLM, vLLM, and SGLang to drive advancements in the field.
Verify that new GPU product launches produce industry leading performance.
Collaborate across the company to guide the direction of inference serving, working with software, research, and product teams to ensure best-in-class performance.
Use the latest coding agents and inference technology to improve team efficiency.

What we need to see:

Master's or PhD degree in Computer Science, Computer Engineering, related fields, or equivalent experience.
Relevant software development experience.
Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
Experience developing client server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.
Solid understanding of CPU and GPU microarchitecture and performance characteristics.
Experience with complex software projects like frameworks, compilers, or operating systems.
Demonstrated proficiency with the latest AI coding agents like Claude Code, Codex, and Cursor
Excellent written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.

Ways to stand out from the crowd:

Demonstrate a drive to continuously improve software and hardware performance.
Showcase examples of novel use cases for agentic AI tools in the workplace.
Experience with databases and visualization tools will set you apart.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have a team of highly skilled and motivated individuals who excel in their work. If you have a proactive and independent approach, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD for Level 2, and 152,000 USD - 241,500 USD for Level 3.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until April 26, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

2701 San Tomas Expressway, Santa Clara, CA, United States, Santa Clara

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

NVIDIA

Deep Learning Architect, LLM Inference - New College Grad 2026

Similar Jobs

Compliance Advisory Officer

Compliance Advisory Officer

Retail Contingent

NVIDIA Santa Clara, California, USA Office

What you need to know about the San Francisco Tech Scene

Key Facts About San Francisco Tech