ifm Logo

ifm

Inference Optimization Intern – Performance Modeling

Reposted Yesterday
Be an Early Applicant
In-Office
Sunnyvale, CA, USA
Internship
In-Office
Sunnyvale, CA, USA
Internship
Intern will build and validate a simulator and profiling framework to model inference on NVIDIA GPUs, develop analytical GPU kernel performance models, profile kernels with Nsight tools, analyze PTX/SASS, identify bottlenecks, and recommend kernel and system-level optimizations for transformer inference on Hopper and Blackwell architectures.
The summary above was generated by AI
About the Institute of Foundation Models
 
The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.

Key Responsibilities

    This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.
    Responsibilities include:
  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Compare measured kernel performance against architectural peak throughput.
  • Identify performance bottlenecks in compute, memory, communication, and scheduling.
  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
  • Investigate PTX and SASS code generation to understand low-level execution behavior.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.

Academic Qualifications

    Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.

Preferred Qualifications

  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low-level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Strong programming skills in C++, CUDA, and Python.

Desired Skills

  • Performance engineering mindset.
  • Strong analytical and debugging abilities.
  • Interest in AI systems, inference optimization, and hardware-software co-design.
  • Ability to work independently on research and engineering challenges.
  • Excellent written and verbal communication skills.

Similar Jobs

29 Minutes Ago
In-Office
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead a multidisciplinary software test and automation team to define verification and validation strategies across simulation, SIL/HIL/VIL and operational environments. Build automated regression tests, validation frameworks, and CI/CD-integrated test infrastructure; establish software readiness criteria and metrics; coordinate with engineering, integration, and flight test teams; develop roadmaps and staffing plans; mentor engineers and present readiness and risk to leadership.
Top Skills: Ci/CdHilRosSilSonarqubeVil
An Hour Ago
In-Office or Remote
United States
177K-294K Annually
Senior level
177K-294K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead site management and monitoring across assigned countries/clusters/regions, ensuring site start-up, quality, regulatory and GCP compliance. Line-manage SCP/CRA/monitors, oversee FSP/CRO delivery, build investigator and stakeholder relationships, drive strategic initiatives (e.g., DCT readiness, virtual monitoring), manage resources and risks, and provide regional insights to optimize clinical trial conduct and timelines.
An Hour Ago
Remote or Hybrid
United States
100K-180K Annually
Senior level
100K-180K Annually
Senior level
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Lead intake, roadmap, and governance for an AI automation/platform team. Harden and ship business-built AI apps to production with secure GCP platforms, define acceptance criteria and responsible-AI gates, manage vendor buy-vs-build decisions, drive adoption, measure outcomes, and manage AI economics and stakeholder communication.
Top Skills: AgentforceAuth/IdentityAWSAzureCi/CdDpaFinopsGCPIso 27001LlmsMlopsMonitoringNetSuiteSalesforceSecrets ManagementSlackSoc 2SQL

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account