Tag: #reasoning-models
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
Reasoning Models in 2026 — A Deep Dive on o3, o4, DeepSeek R1, Claude Thinking, Gemini Deep Think, and QwQ
It has been about a year and a half since o1 (Sept 2024) opened the test-time compute axis. In 2026, 'reasoning models' are no longer a separate family — they are a mode that every frontier model can enter. This guide la
2026-05-14 · 20 min read #reasoning-models#o3#o4#deepseek-r1#claude-thinkingOpenAI RFT with Custom Graders: A Practical Guide for Product and Platform Teams
A practical guide to OpenAI reinforcement fine-tuning with custom graders, including when to use it, how to prepare data, how to evaluate checkpoints, and how to roll it out safely.
2026-04-12 · 5 min read #ai-platform#openai#rft#reinforcement-fine-tuning#custom-graders