Tag: #swe-bench
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3 posts
Separating Signal from Noise When You Evaluate AI Coding Models — Why SWE-bench Got Shaky
OpenAI's evals team argues that SWE-bench Verified, the most widely used coding benchmark, no longer gives meaningful signal because of contamination and design flaws. Look at benchmarks through two axes — signal (the po
2026-07-11 · 6 min read #llm-evaluation#coding-benchmarks#swe-bench#benchmarks#ai-codingCode Becomes the Agents Execution Substrate: A Code-as-Harness View
This post reframes code not as an LLMs final artifact but as the execution harness through which an agent interacts with its environment. Centered on verification loops, tool calling, and execution feedback, it lays out
2026-06-25 · 19 min read #ai#agentic-coding#llm-agents#tool-calling#reactAI Agent & LLM Benchmarks 2026 — SWE-bench Verified / ARC-AGI 2 / GAIA / MMLU-Pro / GPQA / LiveCodeBench / Chatbot Arena Deep Dive
A single-page map of the 30+ AI benchmarks that matter in 2026. From SWE-bench / SWE-bench Verified / SWE-bench Multimodal to AgentBench, WebArena and GAIA, ARC-AGI 2 (Chollet $1M prize), RE-Bench (METR), Frontier Math (
2026-05-16 · 19 min read #ai-benchmark#llm-evaluation#swe-bench#swe-bench-verified#agentbench