Tag: #nccl
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 4 posts
torchcomms in PyTorch 2.13 — The New Distributed Communication Backend Coming for c10d
Among the highlights of PyTorch 2.13, released on July 8, 2026, the most structural change is torchcomms — PyTorch Distributed's new communication backend has begun landing in core's CI and DeviceMesh paths. torchcomms i
2026-07-17 · 10 min read #pytorch#distributed-training#nccl#deep-learningDistributed Training & GPU Infrastructure 2026 Deep-Dive — DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, Blackwell GB200, MI325X, TPU v5p
A comparison of DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, and Composer — plus NVIDIA Blackwell GB200 NVL72, AMD MI325X, Intel Gaudi 3, AWS Trainium 2, and Google TPU v5p/v6e Trillium. 3D parallelism, ZeR
2026-05-16 · 14 min read #distributed-training#deepspeed#fsdp#megatron-lm#rayAdvanced CUDA GPU Programming: Warp Optimization, Tensor Cores, and Triton Kernels
A comprehensive deep dive into CUDA memory hierarchy, Warp optimization, Tensor Core WMMA API, Flash Attention implementation, and Triton custom kernel authoring for accelerating AI model training.
2026-03-17 · 19 min read #cuda#gpuprogramming#tensorcore#triton#flash-attentionDistributed Systems Complete Guide: From CAP Theorem to Distributed ML Training and Kafka
A comprehensive guide to distributed systems for AI engineers, covering CAP theorem, Raft consensus, Kafka message queues, Ring-AllReduce distributed ML training, and NCCL.
2026-03-17 · 17 min read #distributed-systems#kafka#raft#pytorchdistributed#nccl