SemiAnalysis Logo

SemiAnalysis

Research Analyst – AI Hardware & System Performance

Reposted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Build first-principles performance and cost models for AI training and inference across chips, accelerators, memory, interconnects, racks, and datacenters. Evaluate architectures, validate models against benchmarks, analyze semiconductor manufacturing and packaging impacts on performance and cost, publish research, and advise clients and investors on hardware trade-offs and TCO metrics.
The summary above was generated by AI

Employment Type: Full-Time
Work Setting: Remote
Work Location: Korea/International
Work Hours: Office hours
Find out more here: https://semianalysis.com

About SemiAnalysis

SemiAnalysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to cutting-edge AI Models, software, and infrastructure. We are recognized as the leading authority on the semiconductor supply chain, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies.

We’re a global team of over 50 analysts, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually.

Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies.

We also offer three core products:

  • Industry Models – we develop and publish industry models on accelerator shipments, datacenter demand and supply, GPU total cost of ownership, and more. We work with hyperscalers, neoclouds, many of the world’s largest hedge funds, and government agencies.

  • Core Research – our public equity markets product, geared towards financial investors, distills our deep technical research and knowledge into key insights on technology and product trends.

  • Consulting and Technical Due Diligence – We conduct custom research and project work to guide key strategic and investment decisions for the largest private‑equity funds, leading venture‑capital firms, companies across the AI ecosystem, and government agencies.

1) Position Overview

We are looking for someone with a deep and genuine obsession with hardware, regardless of where that expertise comes from.

Candidates may have experience in chip architecture, SoC design, memory systems, interconnect PHYs, semiconductor manufacturing, advanced packaging, performance engineering, ML systems, HPC, datacenter infrastructure, networking, power, cooling, or large-scale cluster operations.

What matters most is a first-principles understanding of how modern AI hardware works, from silicon and packaging through racks, networks, and full datacenter systems. The ideal candidate is intensely curious about why certain designs outperform others and can translate technical differences into measurable performance and economic outcomes.

The role focuses on quantitative analysis of AI hardware and systems, including:

  • Hardware and system performance

  • Architecture evaluation

  • Silicon and system cost analysis

  • Power and cooling constraints

  • Total cost of ownership

  • AI infrastructure economics

  • Technology assessment across the full hardware stack

The work will support SemiAnalysis products and research, including the Inference Simulator, InferenceX, Tokenomics Model, Accelerator & HBM Model, and AI Cloud TCO research.

The scope of the role will flex toward the candidate’s strengths, with strong performers helping define their own coverage areas.

The central question behind the role is:

For a given model, latency target, and system configuration, how many tokens per second does each chip and architecture deliver, and what does each token cost?

This is a remote position. Seoul, Korea, or the broader APAC region is preferred for proximity to industry contacts, although strong candidates from any location will be considered.

2) Responsibilities
  • Build first-principles performance models for LLM inference and training.

  • Analyze arithmetic intensity, roofline performance, prefill and decode behavior, KV cache capacity, memory bandwidth requirements, batching dynamics, parallelism strategies, and latency-throughput trade-offs.

  • Evaluate AI hardware across the full technology stack, including GPUs, TPUs, custom ASICs, emerging accelerators, HBM, on-chip SRAM, memory tiering, scale-up fabrics, scale-out networks, racks, pods, and datacenters.

  • Assess architecture and system design trade-offs, including compute versus memory bandwidth, interconnect topology, power availability, cooling requirements, and workload suitability.

  • Connect semiconductor manufacturing decisions to hardware performance, availability, and cost.

  • Analyze process-node choices, die size, reticle limits, advanced packaging, 2.5D integration, 3D stacking, HBM integration, yield, and silicon economics.

  • Examine new chip, system, rack, and cluster announcements using specifications, die shots, rack layouts, benchmarks, and network diagrams.

  • Develop independent and defensible views on real-world performance compared with vendor claims.

  • Validate internal models against published benchmarks and independently gathered performance data.

  • Investigate and explain differences between theoretical peak performance and delivered application performance.

  • Translate technical analysis into economic metrics, including tokens per second per watt, per dollar of capital expenditure, and per megawatt.

  • Contribute to TCO, tokenomics, accelerator, memory, networking, and datacenter research.

  • Publish detailed research on AI accelerators, systems, infrastructure, and performance.

  • Act as a technical authority during client calls, briefings, and discussions with engineering and investment audiences.

  • Collaborate with accelerator, memory, networking, semiconductor, and datacenter analysts to connect silicon-level findings with broader system and industry conclusions.

3) Requirements
  • Deep technical understanding of modern compute hardware at one or more levels of the stack, including silicon, systems, networking, clusters, or datacenters.

  • Strong interest in learning unfamiliar parts of the hardware and infrastructure stack.

  • Strong quantitative reasoning skills and the ability to work from first principles using FLOPs, bytes, bandwidth, latency, joules, watts, and dollars.

  • Ability to build structured models that connect hardware design, workload performance, infrastructure requirements, and economic outcomes.

  • Working knowledge of how AI workloads stress hardware, including memory bandwidth, memory capacity, interconnect communication, collective operations, and real-world compute utilization.

  • Working understanding of semiconductor manufacturing and economics, including process-node trade-offs, die size, yield, die cost, packaging, and HBM integration.

  • Ability to explain why an architecture was designed in a particular way and how those decisions affect manufacturing cost, performance, power, and scalability.

  • Strong written communication skills, with the ability to explain complex technical findings clearly to both engineering and investor audiences.

  • Ability to take an ambiguous technical question from initial definition through analysis, validation, and publication with minimal oversight.

  • Self-driven, intellectually rigorous, detail-oriented, and comfortable challenging assumptions.

  • Genuine enthusiasm for computer hardware, semiconductor technology, AI systems, and infrastructure research.

4) Preferred Skills
  • Hands-on experience with AI training or inference frameworks such as vLLM, SGLang, TensorRT-LLM, PyTorch, or JAX.

  • Experience with GPU programming or kernel development using CUDA, Triton, or HIP.

  • Understanding of transformer inference mathematics, including per-token FLOPs, memory traffic, KV cache sizing, MoE routing, and quantization effects.

  • Experience designing, deploying, benchmarking, or operating large GPU or accelerator clusters.

  • Knowledge of datacenter networking, rack-scale architecture, power distribution, cooling, or large AI infrastructure deployments.

  • Familiarity with training performance metrics and concepts such as model FLOPs utilization, parallelism scaling efficiency, communication overhead, and failure recovery.

  • Experience benchmarking heterogeneous hardware, including NVIDIA GPUs, AMD GPUs, TPUs, custom accelerators, or emerging non-NVIDIA platforms.

  • Hands-on semiconductor process, foundry, product engineering, yield, packaging, or test experience.

  • Experience with die cost modeling, semiconductor manufacturing economics, advanced packaging, or teardown analysis.

  • Ability to interpret die shots, packaging layouts, physical specifications, and teardown data to infer architecture and process decisions.

  • Prior technical publications, research papers, open-source contributions, benchmarks, hardware teardowns, or industry analysis.

  • Experience communicating technical findings directly to clients, executives, engineers, or institutional investors.

5) Growth Areas

The role provides opportunities to expand beyond an existing area of specialization and develop expertise across the complete AI hardware stack.

Potential growth areas include:

  • Expanding from chip-level expertise into rack-scale and datacenter system analysis.

  • Developing deeper knowledge of AI inference and training workloads.

  • Learning first-principles performance and cost modeling.

  • Building expertise in accelerator, HBM, networking, power, cooling, and TCO analysis.

  • Connecting semiconductor manufacturing and packaging decisions to system-level performance and economics.

  • Developing the ability to evaluate new chips, systems, and infrastructure announcements rapidly and independently.

  • Strengthening technical writing and publishing skills.

  • Becoming a recognized technical authority with engineering, investment, and industry audiences.

  • Leading client briefings and translating complex hardware analysis into strategic and financial conclusions.

  • Collaborating across accelerator, memory, networking, semiconductor manufacturing, and datacenter research areas.

  • Defining and owning a research coverage area based on individual technical strengths.

  • Contributing directly to the development of SemiAnalysis models, simulators, datasets, and research methodologies.

Similar Jobs

An Hour Ago
Easy Apply
Remote
United States
Easy Apply
105K-120K Annually
Senior level
105K-120K Annually
Senior level
Cloud • Information Technology • Security • Software
Own end-to-end complex Professional Services engagements: discovery, architecture, and advanced Risk Cloud configuration. Provide strategic GRC advisory, leverage AI-assisted workflows, manage project/account health, mentor junior consultants, and collaborate cross-functionally to deliver measurable customer value and ROI.
Top Skills: AIApi IntegrationsAuditboardDrataLogicgate Risk CloudMetricstreamOnetrustReporting/Analytics Tools
An Hour Ago
Remote or Hybrid
South Coast, CA, USA
15-22 Hourly
Junior
15-22 Hourly
Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Provide friendly, efficient customer service at cash wrap and on the sales floor. Operate POS, process shipments, manage stock levels, execute visual merchandising, support sales activities, maintain housekeeping and loss-prevention standards, and handle physical tasks like lifting and stockroom organization.
Top Skills: Cash Register SystemsInternetIpadMobile PosPosWalkie Talkie
An Hour Ago
Remote or Hybrid
Palo Alto, CA, USA
100K-176K Annually
Mid level
100K-176K Annually
Mid level
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Design, implement, and operate critical, scalable backend services (identity, friend graph, storage). Collaborate across teams, evaluate trade-offs, test and debug, manage availability, scalability, operational excellence, cost, and participate in incident resolution.
Top Skills: AWSC++GCPJavaKotlinKubernetesMemcacheNoSQLPythonRedisSwift

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account