Tag: #model-architecture
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 1 posts
Mixture of Experts (MoE) Architecture Deep Analysis: Evolution from Switch Transformer to Mixtral and Efficient Scaling Strategies
Deep analysis from the core principles of MoE architecture to Switch Transformer single expert routing, Mixtral 8x7B Sparse MoE, and DeepSeek-MoE fine-grained strategy. Covers routing mechanisms, load balancing loss, tra
2026-03-11 · 17 min read #ai-papers#moe#switch-transformer#mixtral#model-architecture