태그: #scaling-laws
GPU·LLM·MLOps·쿠버네티스, 그리고 마음가짐에 관한 글 · 3 편
LLM 랜드마크 논문 2026 완벽 가이드 - Transformer · Scaling Laws · Flash Attention · Mamba · DeepSeek-R1 · Titans 심층 분석
2017년 Attention Is All You Need에서 2026년 Titans와 DeepSeek-R1까지, LLM 시대를 만든 50여 편의 랜드마크 논문을 테마별로 정리한다. Transformer · BERT · GPT 시리즈 · Scaling Laws · Chinchilla · InstructGPT · PaLM · Flash Attention 1/2/3 · LLaMA 1/2/3/4 ·
2026-05-16 · 34 분 읽기 #llm-papers#transformer#scaling-laws#flash-attention#mambaLLM 사전 학습 & 스케일링 법칙: Chinchilla, Flash Attention, MoE까지
Chinchilla 스케일링 법칙, Common Crawl 데이터 준비, Flash Attention 2, GQA, MoE 아키텍처부터 DeepSeek-V3, Llama 3.1 사전 학습 레시피까지 LLM 사전 학습 완전 가이드입니다.
2026-03-17 · 22 분 읽기 #llm-pretraining#scaling-laws#chinchilla#flash-attention#mixtralmoe대규모 모델 학습 완전 가이드: 100B+ 파라미터 LLM 사전학습 전략
수백억 파라미터 LLM을 실제로 학습시키는 전략과 기법 완전 가이드. 스케일링 법칙(Chinchilla), Megatron-LM, 3D 병렬화, 체크포인팅 전략, 학습 안정성, 데이터 혼합 전략까지 실전으로 배웁니다.
2026-03-17 · 29 분 읽기 #large-scale-training#llm#megatron-lm#distributed-training#scaling-laws