Tag: #red-team
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
AI Safety & Alignment 2026 Deep Dive - Constitutional AI · RLHF · DPO · GRPO · Mechanistic Interpretability · AISI Evals · Red Team
A single-shot map of AI safety and alignment as of 2026. Starts from conceptual roots like outer/inner alignment and mesa-optimization, walks through training-time alignment (RLHF, DPO, GRPO, Constitutional AI), frontier
2026-05-16 · 20 min read #ai-safety#ai-alignment#constitutional-ai#rlhf#dpoThe Complete Guide to LLM Security: Prompt Injection, Jailbreak, Red Team, OWASP LLM Top 10, EU AI Act (2025)
Across 2024–2025 a "security incident" is always among the TOP 3 causes of LLM product failure. 12 prompt injection variants, jailbreak techniques, data exfiltration, model extraction, red team automation (PyRIT/Garak),
2026-04-15 · 12 min read #llm-security#prompt-injection#jailbreak#red-team#owasp-llm