Tag: #tpu
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 5 posts
Systolic Arrays and Dataflow Architecture — The Heart of the TPU
A deep dive into the systolic array, the structure that lets AI accelerators run matrix multiplication efficiently, complete with ASCII diagrams. We walk through dataflow strategies like weight-stationary and output-stat
2026-06-16 · 20 min read #gpu-cuda#systolic-array#dataflow#tpu#ai-hardwareGPU vs TPU vs ASIC — The 2026 Inference War
A comparison of the GPU, TPU, and ASIC competition over 2026 inference workloads. We cover Google TPU v6 Trillium and Ironwood, the fast-growing cloud in-house inference ASICs, the throughput/latency/cost/power trade-off
2026-06-16 · 23 min read #gpu#tpu#asic#inference#ai-hardwareDistributed Training & GPU Infrastructure 2026 Deep-Dive — DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, Blackwell GB200, MI325X, TPU v5p
A comparison of DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, and Composer — plus NVIDIA Blackwell GB200 NVL72, AMD MI325X, Intel Gaudi 3, AWS Trainium 2, and Google TPU v5p/v6e Trillium. 3D parallelism, ZeR
2026-05-16 · 14 min read #distributed-training#deepspeed#fsdp#megatron-lm#rayGoogle TPU Deep Dive: How Systolic Arrays Solve Matrix Multiplication Perfectly
A complete technical breakdown of how Google's Systolic Array achieves extreme efficiency for matrix multiplication. From INT8 inference and bfloat16, to XLA compiler optimizations and TPU Pod distributed inference - wit
2026-03-18 · 15 min read #tpu#google#systolic-array#model-serving#jaxAI Hardware Accelerators Complete Guide: H100, TPU, Cerebras, and Edge AI Chips Compared
A comprehensive comparison guide covering NVIDIA H100 Tensor Core, Google TPU v5 systolic array, Cerebras WSE-3, AWS Inferentia 2, and Apple Neural Engine for AI hardware accelerators.
2026-03-17 · 17 min read #ai-hardware#h100#tpu#cerebras#edge-ai