Tag: #inference-cost
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3 posts
The Conditions Under Which a 9B Fine-Tuned for 500 Dollars Beat the Frontier — And How Narrow They Are
On July 28, 2026, a post scored 336 points on Hacker News. It reports that Fermisense trained Qwen3.5-9B with GRPO on roughly 500 dollars worth of GPU time and beat five frontier configurations — using the same tools and
2026-07-31 · 11 min read #ai#llm#fine-tuning#reinforcement-learning#inference-costGPT-5.6 and the Limits of Price-Performance — How to Find Your Workload's Place on the Curve
On July 30, 2026, OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. That is three weeks after the July 9 general availability, and the top-end Sol price is unchanged. Luna went from 1 dollar input and 6 dollars output per
2026-07-31 · 12 min read #ai#llm#openai#inference-cost#prompt-cachingWhat DeepSeek V4 Flash Actually Changes — A Cost-per-Capability Question, Not a Leaderboard One
On July 31, 2026, DeepSeek moved its V4-Flash API into public beta. The architecture is identical to April's preview — 284B total parameters, 13B active MoE, 1M context — and the only thing that changed is post-training.
2026-07-31 · 11 min read #ai#llm#deepseek#inference-cost#moe