Tag: #red-teaming
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3 posts
AI Safety, Evals and Red-Teaming in 2026 — Deep Dive into Inspect AI, Garak, PyRIT, Promptfoo, OpenAI Evals, lm-eval-harness
A single-page map of the 2026 AI safety, evaluation, and red-teaming ecosystem. Inspect AI (Anthropic, adopted by UK AISI), Garak (NVIDIA then independent), PyRIT (Microsoft), Promptfoo (YC), OpenAI Evals, lm-evaluation-
2026-05-16 · 22 min read #ai-safety#red-teaming#evaluation#inspect-ai#garakAI Safety & Alignment Complete Guide 2025: Responsible AI, RLHF, Constitutional AI, Red Teaming
Everything about AI Safety! Alignment problem (goal alignment), RLHF/DPO/Constitutional AI, Bias detection/mitigation, Hallucination prevention, Red team testing, AI Guardrails, Interpretability (SHAP/LIME), EU AI Act, E
2026-04-14 · 25 min read #ai-safety#alignment#responsible-ai#rlhf#constitutional-aiLLM Safety and Red Teaming Practical Guide: From Adversarial Defense to Guardrail Implementation
A practical guide to LLM safety covering red teaming methodology, adversarial attack defense, and guardrail implementation.
2026-03-08 · 42 min read #llm#red-teaming#safety#guardrails#prompt-injection