タグ: #torchtitan
GPU・LLM・MLOps・Kubernetes、そしてマインドセット · 1 件
分散学習 & GPUインフラ 2026 ディープダイブ — DeepSpeed、FSDP2、Megatron-LM、Ray Train、JAX、TorchTitan、Blackwell GB200、MI325X、TPU v5p 総まとめ
DeepSpeed/FSDP2/Megatron-LM/Ray Train/JAX/TorchTitan/Composerを比較し、NVIDIA Blackwell GB200 NVL72、AMD MI325X、Intel Gaudi 3、AWS Trainium 2、Google TPU v5p/v6e Trillimuまで。3D並列化、ZeRO-FSDP等価性、MoEのAll-to-All、fp8/mxfp4精度、NCCLチューニン
2026-05-16 · 20 分で読めます #distributed-training#deepspeed#fsdp#megatron-lm#ray