Tag: #deepseek-r1
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
LLM Landmark Papers Roundup 2026 - Transformer / Scaling Laws / Flash Attention / Mamba / DeepSeek-R1 / Titans Deep Dive
From the 2017 Attention Is All You Need paper to 2026 Titans and DeepSeek-R1, a thematic roundup of the 50+ landmark papers that built the LLM era. Transformer, BERT, GPT 1-3, Scaling Laws, Chinchilla, InstructGPT, PaLM,
2026-05-16 · 21 min read #llm-papers#transformer#scaling-laws#flash-attention#mambaReasoning Models in 2026 — A Deep Dive on o3, o4, DeepSeek R1, Claude Thinking, Gemini Deep Think, and QwQ
It has been about a year and a half since o1 (Sept 2024) opened the test-time compute axis. In 2026, 'reasoning models' are no longer a separate family — they are a mode that every frontier model can enter. This guide la
2026-05-14 · 20 min read #reasoning-models#o3#o4#deepseek-r1#claude-thinking