Tag: #mixtralmoe
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 1 posts
LLM Pretraining & Scaling Laws: From Chinchilla to Flash Attention and MoE
A complete guide to LLM pretraining: Chinchilla scaling laws, Common Crawl data pipelines, Flash Attention 2, GQA, MoE architectures, and the latest pretraining recipes from DeepSeek-V3, Llama 3.1, and phi-4.
2026-03-17 · 14 min read #llmpretraining#scalinglaws#chinchilla#flash-attention#mixtralmoe