Tag: #qwen
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 6 posts
Open-Source LLMs 2026 Deep Dive - Llama 4 · DeepSeek V3 + R1 · Qwen 3 · Mistral Large 2 · Phi-4 · Gemma 3 · Falcon 3
In spring 2026, open-source LLMs are no longer the shadow of closed models. Meta Llama 4 (Scout 109B, Maverick 400B MoE, Behemoth 2T), Llama 3.3 70B as the last dense baseline, DeepSeek V3 671B MoE and the R1 reasoning m
2026-05-16 · 32 min read #open-source-llm#llama-4#deepseek#qwen#mistralTop LLM Papers 2024-2026 - Llama, DeepSeek, Qwen, Mistral, Phi, RLHF, DPO, CoT, RAG, FlashAttention, vLLM Reading List
A curated reading list of 30+ must-read LLM papers for engineers building with LLMs in 2024-2026. Covers foundation models (Llama 3/4, DeepSeek-V3/R1, Qwen3, Mistral, Phi-4, Gemma 3), training innovations (MoE, MLA, GQA)
2026-05-16 · 19 min read #llm#papers#llama#deepseek#qwenChinese AI Labs in 2026 — A Deep Dive into DeepSeek, Qwen, Kimi, GLM, Yi, Doubao, Hunyuan (The New Center of Gravity for Open Weights)
When DeepSeek-V3 dropped a 671B MoE in December 2024 and R1 added open-weight reasoning in January 2025, the world paused. Since then Alibaba's Qwen 3 (235B-A22B) has become the de facto new open-weight standard; Moonsho
2026-05-15 · 25 min read #ai#china#deepseek#qwen#alibaba2025 Open Source AI Models Showdown: DeepSeek R1 vs Llama 4 vs Qwen 3 vs Mistral
DeepSeek R1 (671B/37B), Llama 4 Scout/Maverick, Qwen 3 (235B MoE), Mistral 8x22B — complete comparison of the 2025 open-source AI model leaders with benchmarks, licenses, deployment guides, and cost analysis.
2026-03-22 · 20 min read #open-source#ai#llm#deepseek#llamaComplete Guide to Open Source LLMs: Llama 3, Mistral, DeepSeek, Qwen, and Gemma
A comprehensive overview of the open source LLM landscape covering Llama 3, Mistral, DeepSeek, Qwen, and Gemma.
2026-03-17 · 14 min read #llm#llama#mistral#deepseek#qwenOpen-Source LLM Landscape Guide: Models, Tools, and Deployment in 2026
A comprehensive guide to the open-source LLM ecosystem in 2026. Covers the leading model families (Llama, Mistral, Gemma, Qwen, DeepSeek), local inference tools (Ollama, llama.cpp, vLLM), fine-tuning techniques (LoRA, QL
2026-03-17 · 20 min read #open-source#llm#llama#mistral#gemma