태그: #megatron-lm
GPU·LLM·MLOps·쿠버네티스, 그리고 마음가짐에 관한 글 · 2 편
분산 학습 & GPU 인프라 2026 딥다이브 — DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, Blackwell GB200, MI325X, TPU v5p 총정리
DeepSpeed/FSDP2/Megatron-LM/Ray Train/JAX/TorchTitan/Composer를 비교하고, NVIDIA Blackwell GB200 NVL72, AMD MI325X, Intel Gaudi 3, AWS Trainium 2, Google TPU v5p/v6e Trillium까지. 3D 병렬화, ZeRO-FSDP 등가성, MoE All-to-All, fp8/mxfp
2026-05-16 · 23 분 읽기 #distributed-training#deepspeed#fsdp#megatron-lm#ray대규모 모델 학습 완전 가이드: 100B+ 파라미터 LLM 사전학습 전략
수백억 파라미터 LLM을 실제로 학습시키는 전략과 기법 완전 가이드. 스케일링 법칙(Chinchilla), Megatron-LM, 3D 병렬화, 체크포인팅 전략, 학습 안정성, 데이터 혼합 전략까지 실전으로 배웁니다.
2026-03-17 · 29 분 읽기 #large-scale-training#llm#megatron-lm#distributed-training#scaling-laws