LabHub

Blog

AI Hardware Accelerators 2026 — NVIDIA Blackwell / AMD Instinct MI400 / Google TPU Trillium / Cerebras WSE-3 / Groq LPU / Tenstorrent / Etched Sohu / Furiosa / Rebellions Deep Dive

한국어English日本語

1. The 2026 AI Hardware Map — Four Camps: Hyperscaler / Challenger / In-House / Edge

In May 2026 the AI chip market looks nothing like it did five years ago. The single-vendor era of NVIDIA — V100 in 2020, A100 in 2021, H100 in 2022, H200 in 2023 — opened a new chapter when Blackwell debuted at GTC 2024. And in 2026, there are more chips than ever, and the choice has only gotten harder.

Roughly four camps now exist.

On pricing: an H100 was 30K40Kin2024;B200landsat30K-40K in 2024; B200 lands at 30K-40K as well, and a GB200 NVL72 rack runs about $3M. Cloud rentals settled at $2-4 per hour for H100 and now $4-8 per hour for partial B200 instances.

This piece walks through each camp — specs, architecture, memory and interconnect, and finally the Korean and Japanese players.

All numbers are sourced from public material as of May 2026, plus reporting from SemiAnalysis, The Information and Reuters. Private cluster prices are estimates.


2. NVIDIA Blackwell — B100 / B200 / GB200 NVL72 / B300 / Rubin

Family structure

Blackwell is NVIDIA's fifth generation data-center GPU architecture, unveiled by Jensen Huang at GTC March 2024. It is the successor to Hopper (H100/H200), fabbed on TSMC N4P, and for the first time uses a chiplet structure — two GPU dies connected by NV-HBI (NVIDIA High-Bandwidth Interconnect) at 10 TB/s.

Why NVL72 matters

Seventy-two B200 GPUs share one NVLink domain. A model sees 72 GPUs as if they were one. MoE token routing — the all-to-all step — happens inside NVLink, never on InfiniBand. That removes the real bottleneck behind GPT-4 / Claude 3.5 class training.

Rubin — September 2026

Rubin is NVIDIA's sixth generation. Pre-announced at GTC 2024, formal launch is expected at GTC September 2026.

NVIDIA's annual cadence holds: 2024 Blackwell, 2025 Blackwell Ultra, 2026 Rubin, 2027 Rubin Ultra.

Pricing and supply

An H100 was $30K-40K per card in 2024. B200 trades at $30K-40K per card and $4-8 per hour in cloud partials. A GB200 NVL72 rack runs about $3M. In 1H 2025 NVIDIA shipped more than two million Blackwell GPUs per quarter (Reuters).


3. AMD Instinct — MI300X to MI325X to MI355X to MI400 Helios

MI300X (December 2023)

CDNA 3 architecture, 192GB HBM3, 5.2 PFLOPS FP8. With 2.4x H100's memory (80GB), it landed at Meta and Microsoft for Llama inference fleets. About $15K-20K per card.

MI325X (Q4 2024)

Memory bumped to 256GB HBM3E, clocks raised slightly. The H200 counter.

MI355X (late 2025)

CDNA 4 architecture. 288GB HBM3E, native FP4 datatype. The direct response to Blackwell B200/B300. ROCm 6.x has matured to where PyTorch / vLLM / SGLang feel almost as smooth as on NVIDIA.

MI400 Helios (2026)

The next-generation platform AMD unveiled at Advancing AI 2025.

UALink is the open alternative to NVLink. AMD / Broadcom / Cisco / Google / HPE / Intel / Meta / Microsoft formed the consortium, and the 1.0 spec was published in 1H 2026.

Market share

In 2025 data-center GPU revenue, NVIDIA held 90%+, AMD 5-7%, with Intel and in-house chips splitting the remainder. AMD locked in Microsoft Azure ND-MI355X-v6 and a Meta cluster on MI355X, and Oracle Cloud announced the first big MI400 Helios deployment.


4. Intel Gaudi 3 + Falcon Shores Rumors

Gaudi 3 — the last standalone line

Intel acquired Habana Labs for about $2B in 2019 and has shipped Gaudi 1 / 2 / 3 since. Gaudi 3 launched April 2024 on TSMC N5, with 128GB HBM2E and 8x 200 Gbps Ethernet (RoCE instead of InfiniBand) interconnect.

Stability AI, Naver and Intel's own Tiber Cloud are notable adopters.

Falcon Shores rumors

Falcon Shores was originally the planned Gaudi-plus-Ponte-Vecchio fusion shipping in 2024. In September 2024 Intel officially canceled the external launch, keeping it as an internal R&D vehicle.

As of May 2026, the rumor is that Intel is preparing either a Gaudi 4 or a brand-new GPU line for 1H 2027. The seed of that story is a comment from Lip-Bu Tan (the new CEO since 2025) at an IFS Cup event that Intel will "reorganize the AI chip line."


5. Apple M5 + M5 Pro + Neural Engine + AC1 Server Chip

M5 / M5 Pro / M5 Max — October 2025

Apple Silicon fifth generation, on TSMC N3E. CPU core counts are unchanged. The GPU gains ray-tracing accelerators and a new matrix engine aimed at AI inference.

The Neural Engine stays at 16 cores. The change is matrix-multiply throughput and INT4 quantization acceleration.

AC1 server chip — spring 2026 rumor

Sourced from The Information (November 2025) and Bloomberg's Mark Gurman. Apple is building a server-class SoC for data-center AI inference.

Apple already runs Apple Intelligence Private Cloud Compute (PCC) on M2 Ultra Macs. AC1 is the next-generation PCC silicon.


6. Google TPU v5p + Trillium (v5/v6)

TPU lineage

What Trillium does

Announced at Google I/O May 2024. The workhorse for Gemini 2.0 training.

Trillium scales by pod: 256 chips per pod, with optical ICI (Inter-Chip Interconnect) reaching 8,960 chips in a SuperPod.

TPU gen 7 — late 2026 rumor

The Information reports a TPU v7 reveal slated for late 2026. With Anthropic relying heavily on TPUs, the stakes are large.


7. Cerebras WSE-3 — 4 Trillion Transistors, Wafer Scale

The wafer-scale idea

A standard chip is cut from a 12-inch wafer into reticle-sized dies (about 858 mm²). Cerebras uses the entire wafer as one chip.

WSE-3 (announced March 2024):

Why wafer scale

Eliminate chip-to-chip communication. Memory (SRAM) sits directly next to compute, giving bandwidth orders of magnitude beyond HBM. All model weights live in on-wafer SRAM — a 70B model fits on one wafer.

Tradeoffs

A CS-3 system is estimated at $2-3M.


8. Groq LPU — Sequential Inference Speed

The LPU idea

Groq's LPU (Language Processing Unit) is the chip from a company founded by Jonathan Ross, a former Google TPU engineer, in 2016. Deterministic execution — every instruction runs on a cycle precisely scheduled by the compiler.

Why it's fast

GPUs use dynamic scheduling to spread work across SMs. The LPU resolves all dispatch at compile time — no runtime branching. The result: Llama 70B inference at 200-300 tokens per second, four to eight times faster than the 30-50 tokens per second on an H100.

Tradeoffs

That said, the LPU shines at latency-first workloads: code completion, chatbots, voice assistants. Groq Cloud serves Llama 70B starting at $0.59 per hour.


9. SambaNova SN40L — Reconfigurable Dataflow

SambaNova's bet

Founded in 2017 by Stanford professor Kunle Olukotun and Rodrigo Liang. The Reconfigurable Dataflow Architecture (RDA) reshapes the on-chip data flow per workload.

SN40L (2023):

Why RDA

GPUs are SIMT machines tuned for dense tensor math. Transformers add irregular patterns — dynamic-shape attention, sparse MoE dispatch. RDA compiles a custom data path per layer, making it strong on sparse workloads.

Customers

U.S. DOE labs (Lawrence Livermore, Argonne), Saudi Aramco, parts of the SoftBank R&D cluster.


10. Tenstorrent — Jim Keller, RISC-V Open Architecture

Jim Keller's company

Former AMD Zen lead architect, former Apple A4/A5 lead, former Tesla Autopilot chip lead, former Intel SVP. He joined Tenstorrent as CEO in 2021.

The core differentiators:

Lineup

Hyundai / Samsung / LG AI Research investment

A 2024 Korean consortium (Hyundai Motor, Samsung NEXT, LG) invested in Tenstorrent. Korea sees automotive AI and data-center AI as the targets.


11. Etched Sohu — Transformer-Only ASIC (June 2024)

One thing, done well

Etched was founded in 2022 by two Harvard undergrads. In June 2024 the Sohu unveil made waves.

Why transformer-only

Less than 30% of a GPU's area actually services transformer inference. Since attention and FFN patterns are so well known, strip the other 70% of silicon and pack more attention units in its place.

Risk and reward

The risk is obvious. If Mamba / RWKV / SSM / diffusion architectures rise, Sohu becomes useless overnight. As of May 2026, transformers still account for 80%+ of LLMs — Etched is betting hard on that.

Series A in 2024 was $120M, with Peter Thiel and Stanley Druckenmiller among the investors.


12. AWS Trainium 2 + Inferentia 3

The AWS in-house silicon strategy

AWS shipped Inferentia 1 in 2018, Trainium 1 in 2020, Inferentia 2 in 2023, Trainium 2 in 2024, and Inferentia 3 in 2025.

A Trn2.48xlarge instance has 16 chips and 1.5TB HBM at roughly $5-6 per hour.

Anthropic's Project Rainier

Anthropic announced Project Rainier in 2024 — a massive Trainium 2 cluster. Reports put it at 400,000 Trainium 2 chips, used to train Claude 4.x (official statement).

AWS will launch Trainium 3 in late 2026. The Neuron SDK now feels native in PyTorch and JAX.


13. MatX / Tachyum Prodigy — Newer Entrants

MatX

Founded 2022 by former Google TPU and OpenAI engineers. The mission: a chip purpose-built for LLM training. Series B raised $80M in 2025; the first silicon targets late 2026.

Tachyum Prodigy

Founded by Slovak-origin Radoslav Danilak. The ambition: AI plus HPC plus general compute on a single chip.

The skeptics are loud but EuroHPC (the EU public HPC program) is positioned as the first big customer.


14. Phone NPUs — A18 Pro / Snapdragon 8 Gen 4 / Dimensity 9400 / Tensor G5

Apple A18 Pro (September 2024, iPhone 16 Pro)

Snapdragon 8 Gen 4 (October 2024, Samsung S25 and others)

MediaTek Dimensity 9400 (October 2024)

Google Tensor G5 (October 2024, Pixel 9)

The point of phone NPUs is simple: on-device inference is effectively zero cost. No cloud call — LLM responses are generated locally.


NVLink 6 lands with Rubin (2026), at an estimated 3.6 TB/s per chip.

PCIe Gen 6 / Gen 7

The Gen 6 shift introduces PAM4 signaling — solving SerDes limits with four levels instead of two.

CXL

Compute Express Link. The Intel-led standard for memory sharing across CPU, GPU, DPU and memory pools over PCIe.

As of May 2026, CXL 3.0 volume parts (Samsung CMM-D, Micron CZ120) are deployed. NVMe plus CXL memory expansion is the new paradigm for Tier 1 / Tier 2 / Tier 3 memory hierarchies.

The open alternative to NVLink. AMD / Broadcom / Cisco / Google / HPE / Intel / Meta / Microsoft consortium. The 1.0 spec was published in 2026.


16. Memory — HBM3E / HBM4 / Samsung + SK Hynix

HBM evolution

HBM sits next to the GPU die in a 2.5D or 3D stack. With HBM3E and an 8-stack design, bandwidth crosses 8 TB/s.

Supply — SK Hynix / Samsung / Micron

One Blackwell card carries eight HBM3E stacks for 192GB total. A stack is roughly $250-300, so HBM alone costs about $2-2.4K per chip.

HBM4

JEDEC ratified the standard in April 2025. 16 Gbps per pin, 12-Hi and 16-Hi stacks. Volume debut is in Rubin (2026). Both Korean vendors are racing for NVIDIA's qualification.


17. Korea — FuriosaAI + Rebellions (Sapeon Merger 2024)

FuriosaAI

Founded in 2017 by June Paik, a Samsung and AMD veteran. RNGD (Renegade) launched in 2024, targeting Llama inference workloads.

LG AI Research adopted RNGD for EXAONE inference; a collaboration with Kakao Enterprise Cloud was also announced.

Rebellions + Sapeon merger

In July 2024 Rebellions and Sapeon announced a merger. The post-merger entity is named Rebellions. KT, SK Telecom and Samsung all stayed on as investors. The next-gen REBEL chip was unveiled in 2025 and enters volume in 2026.

Korea's K-Cloud initiative aims to deploy domestic AI accelerators in 50% of NIA data centers by 2030.


18. Japan — SoftBank Graphcore + Preferred Networks MN-3 + Rapidus 2nm 2027

SoftBank's Graphcore acquisition (July 2024)

SoftBank acquired UK-based Graphcore for about $500M. Graphcore's IPU (Intelligence Processing Unit) family — Bow IPU and second-generation Colossus — will be folded into SoftBank's Cristal Intelligence (in-house AI infrastructure).

Preferred Networks MN-3 / MN-Core 2

Preferred Networks is Japan's flagship AI company. The MN-Core line is its proprietary training accelerator.

Used internally to train PFN's own LLMs and shared with select partners (notably Toyota) rather than sold externally.

Rapidus — 2nm in 2027

A new foundry funded by the Japanese government, Sony, Toyota, NTT and SoftBank. The goal is 2nm volume in 2027. The technology partnership is with IBM; a fab is under construction in Chitose, Hokkaido.

The U.S., Korea and Taiwan (TSMC) dominate the leading-edge foundry world — Japan is making another run. A pilot line is running as of May 2026; if the 2027 schedule holds, Rapidus becomes the biggest Japanese AI-chip variable of the decade.


19. Liquid Cooling + Data-Center Power

Why liquid cooling

H100 was 700W, B200 is 1000W, a GB200 NVL72 rack is 120 kW. Air cooling can't handle it. Eight 1000W GPUs in a 1U server is 8 kW per server. About 30 kW per rack is the air-cooling ceiling — above that liquid cooling is mandatory.

Forms of liquid cooling

GB200 NVL72 standardizes on D2C. A facility-wide water loop is mandatory. PUE drops to about 1.05 (vs 1.4 to 1.6 for air).

Power — data centers next to power plants

The new generation of Anthropic, OpenAI and Meta data centers is 2 GW and up. That's the consumption of about 2 million U.S. homes.

As of May 2026, AI sites are spreading across U.S. PJM, Texas ERCOT, Taiwan's Hsinchu, Korea's Anseong and Pyeongtaek, and the wider Tokyo-area grid — and the generation infrastructure has become the real bottleneck.


20. Who Should Pick What — Training / Inference / Edge / Phone

Training — big models, new models

SituationRecommendation
Cutting-edge 70B+ MoE trainingNVIDIA GB200 NVL72 / Rubin (late 2026)
Cost-optimized training (50%+ cheaper)AMD MI355X / MI400 Helios
TPU-friendly (JAX / TF)Google TPU v5p / Trillium
OK with AWS lock-inAWS Trainium 2

Inference — high volume

SituationRecommendation
General LLM servingNVIDIA H200 / B200 / AMD MI300X
Ultra-low latency (code completion)Groq LPU / Cerebras WSE-3
Transformer-onlyEtched Sohu (post-launch)
Korea / EXAONE / domestic modelsFuriosaAI RNGD / Rebellions REBEL

Edge — robots, vehicles, IoT

SituationRecommendation
Autonomous drivingNVIDIA Drive Thor / Tesla FSD HW5
Industrial IoTNVIDIA Jetson Orin / Hailo-10 / Tenstorrent
Desktop workstationNVIDIA RTX 5090 / AMD Radeon Pro

Phone — on-device LLM

SituationRecommendation
iOS Apple IntelligenceA18 Pro Neural Engine
Android Gemini NanoSnapdragon 8 Gen 4 / Tensor G5
Budget AndroidDimensity 9400

The decision criteria are simple — software-stack compatibility, unit cost, and availability. NVIDIA's CUDA ecosystem is still the strongest moat, but ROCm, XLA, Neuron and SynapseAI are closing in.


21. References

Comments

No comments yet.

Sign in to leave a comment