Tag: #safety
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 5 posts
Ferrocene's Certified core — The Compiler Is ASIL D, So Why Does the Library Stop at ASIL B?
The claim that Ferrocene's Rust compiler is 'ASIL D certified' is true, but the conclusion people usually draw from it is not. Ferrocene qualified the compiler as an ISO 26262 ASIL D / TCL 3 tool, which means 'you may de
2026-07-16 · 23 min read #embedded#rust#ferrocene#safety#systemsWhy RLHF Models Game Their Rewards — The Mechanisms, Symptoms, and Mitigations of Reward Hacking
"Reward Hacking in the Era of Large Models," posted to arXiv in April 2026 by Xiaohua Wang and 22 co-authors, is a survey of why and how RLHF-aligned large models game their reward signals. Its central proposal is the Pr
2026-07-11 · 5 min read #ai#llm#alignment#rlhf#safetyRobot Safety and Alignment — Trusting Powerful Robots
How can we trust increasingly powerful robots. We take a balanced look at physical safety, constrained reinforcement learning and safety layers, handling distribution shift, human-robot collaboration safety, verification
2026-06-29 · 17 min read #ai-papers#robotics#safety#alignment#reinforcement-learningComplete Guide to Chatbot Guardrails and Safety: From Prompt Injection Defense to Output Validation
A comprehensive guide to securing production chatbots. Covers prompt injection attack types and defenses, NeMo Guardrails/Guardrails AI frameworks, content filtering, output validation, and PII masking with practical cod
2026-03-13 · 24 min read #chatbot#guardrails#prompt-injection#safety#content-filteringLLM Safety and Red Teaming Practical Guide: From Adversarial Defense to Guardrail Implementation
A practical guide to LLM safety covering red teaming methodology, adversarial attack defense, and guardrail implementation.
2026-03-08 · 42 min read #llm#red-teaming#safety#guardrails#prompt-injection