Tag: #openai-evals
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
AI Safety, Evals and Red-Teaming in 2026 — Deep Dive into Inspect AI, Garak, PyRIT, Promptfoo, OpenAI Evals, lm-eval-harness
A single-page map of the 2026 AI safety, evaluation, and red-teaming ecosystem. Inspect AI (Anthropic, adopted by UK AISI), Garak (NVIDIA then independent), PyRIT (Microsoft), Promptfoo (YC), OpenAI Evals, lm-evaluation-
2026-05-16 · 22 min read #ai-safety#red-teaming#evaluation#inspect-ai#garakAgent Evaluation Systems in 2026 — Inspect AI vs Promptfoo vs Phoenix vs LangSmith vs OpenAI Evals (You're Measuring the Agent, Not the Model)
LLM evals measure the model. Agent evals measure whether the model plus the harness plus the tools actually carry a task to completion. They are different problems. This is a map of the 2026 landscape — Inspect AI from U
2026-05-14 · 20 min read #agent-evaluation#inspect-ai#promptfoo#phoenix#langsmith