Tag: #deepseek
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 13 posts
What DeepSeek V4 Flash Actually Changes — A Cost-per-Capability Question, Not a Leaderboard One
On July 31, 2026, DeepSeek moved its V4-Flash API into public beta. The architecture is identical to April's preview — 284B total parameters, 13B active MoE, 1M context — and the only thing that changed is post-training.
2026-07-31 · 11 min read #ai#llm#deepseek#inference-cost#moeOpen-Source LLMs 2026 Deep Dive - Llama 4 · DeepSeek V3 + R1 · Qwen 3 · Mistral Large 2 · Phi-4 · Gemma 3 · Falcon 3
In spring 2026, open-source LLMs are no longer the shadow of closed models. Meta Llama 4 (Scout 109B, Maverick 400B MoE, Behemoth 2T), Llama 3.3 70B as the last dense baseline, DeepSeek V3 671B MoE and the R1 reasoning m
2026-05-16 · 32 min read #open-source-llm#llama-4#deepseek#qwen#mistralTop LLM Papers 2024-2026 - Llama, DeepSeek, Qwen, Mistral, Phi, RLHF, DPO, CoT, RAG, FlashAttention, vLLM Reading List
A curated reading list of 30+ must-read LLM papers for engineers building with LLMs in 2024-2026. Covers foundation models (Llama 3/4, DeepSeek-V3/R1, Qwen3, Mistral, Phi-4, Gemma 3), training innovations (MoE, MLA, GQA)
2026-05-16 · 19 min read #llm#papers#llama#deepseek#qwenChinese AI Labs in 2026 — A Deep Dive into DeepSeek, Qwen, Kimi, GLM, Yi, Doubao, Hunyuan (The New Center of Gravity for Open Weights)
When DeepSeek-V3 dropped a 671B MoE in December 2024 and R1 added open-weight reasoning in January 2025, the world paused. Since then Alibaba's Qwen 3 (235B-A22B) has become the de facto new open-weight standard; Moonsho
2026-05-15 · 25 min read #ai#china#deepseek#qwen#alibabaLLM Landmark Papers Guide — From Attention to GPT, LLaMA, DeepSeek, o1, and Claude (with References, 2026)
Where do the real shifts in LLMs come from? From Attention is All You Need in 2017 to the reasoning models of 2026, this guide organizes the 20-odd landmark papers you must know, by era and theme. Each paper is compresse
2026-05-14 · 15 min read #llm#research-papers#transformer#gpt#llama2025 Open Source AI Models Showdown: DeepSeek R1 vs Llama 4 vs Qwen 3 vs Mistral
DeepSeek R1 (671B/37B), Llama 4 Scout/Maverick, Qwen 3 (235B MoE), Mistral 8x22B — complete comparison of the 2025 open-source AI model leaders with benchmarks, licenses, deployment guides, and cost analysis.
2026-03-22 · 20 min read #open-source#ai#llm#deepseek#llamaMarch 2025 Tech·AI·K-POP Weekly Digest: From GTC to BTS Comeback
A comprehensive roundup of March 2025 highlights: NVIDIA GTC Blackwell Ultra announcement, Gemini 2.5 Pro topping benchmarks, MCP becoming the industry standard, DeepSeek-R1 open-source shock, BTS full-group comeback aft
2026-03-21 · 12 min read #culture#ai#kpop#nvidia#gtcComplete Guide to Open Source LLMs: Llama 3, Mistral, DeepSeek, Qwen, and Gemma
A comprehensive overview of the open source LLM landscape covering Llama 3, Mistral, DeepSeek, Qwen, and Gemma.
2026-03-17 · 14 min read #llm#llama#mistral#deepseek#qwenLLM Pretraining & Scaling Laws: From Chinchilla to Flash Attention and MoE
A complete guide to LLM pretraining: Chinchilla scaling laws, Common Crawl data pipelines, Flash Attention 2, GQA, MoE architectures, and the latest pretraining recipes from DeepSeek-V3, Llama 3.1, and phi-4.
2026-03-17 · 14 min read #llmpretraining#scalinglaws#chinchilla#flash-attention#mixtralmoeMixture of Experts (MoE) Architecture Paper Deep Analysis: From GShard to DeepSeek-MoE
Analyzes core papers on Mixture of Experts architecture, comparing routing strategies and training stability techniques across GShard, Switch Transformer, Mixtral, and DeepSeek-MoE.
2026-03-14 · 34 min read #ai-papers#moe#transformer#deepseekDeep Dive into Mixture of Experts (MoE) Architecture: From Switch Transformer to Mixtral and DeepSeek
A comprehensive analysis of Mixture of Experts (MoE) architecture. Covers the mathematical foundations of Sparse MoE, routing strategies in Switch Transformer, Mixtral 8x7B, and DeepSeek-V3, training stability techniques
2026-03-10 · 15 min read #ai-papers#moe#transformer#mixtral#deepseekDeep Dive into Sparse Mixture of Experts (MoE) Architecture: From Design Principles to DeepSeek-V3 and Qwen3
Analyzing the mathematical principles, routing strategies, and load balancing techniques of Sparse MoE architecture, covering the design choices and practical training/inference optimization of modern MoE models from Swi
2026-03-06 · 22 min read #ai-papers#moe#sparse-model#deepseek#2026-03Mixture of Experts (MoE) Architecture: A Complete Analysis
A complete analysis of MoE architectures, from the principles of Sparse MoE to the MoE implementations in Mixtral and DeepSeek-V3, routing strategies, and load balancing.
2026-03-03 · 6 min read #ai-papers#moe#mixtral#deepseek#2026-03