Tag: #reliability
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 26 posts
Chaos Engineering Complete Guide 2025: Resilience Testing, Chaos Monkey, Litmus, Game Day
Everything about Chaos Engineering! Chaos principles (hypothesis → experiment → observe → improve), Chaos Monkey/Litmus Chaos/Chaos Mesh, fault injection (network/CPU/memory/Pod/AZ), Game Day operations, progressive adop
2026-04-14 · 21 min read #chaos-engineering#resilience#chaos-monkey#litmus#fault-injectionSRE Practices Guide 2025: Incident Management, Postmortem, Error Budget, On-Call, Toil Elimination
Everything about SRE in practice! Incident management (detect → respond → recover → postmortem), Error Budget policy, On-Call operations (rotation/escalation/fatigue management), Toil elimination automation, SLO/SLI/SLA
2026-04-14 · 26 min read #sre#site-reliability#incident-management#postmortem#error-budgetCloudflare AI Gateway Practical Guide: Observability, Reliability, and Cost Control for AI Traffic
A practical, current guide to Cloudflare AI Gateway as of April 12, 2026, covering observability, caching, retries, rate limiting, model fallback, and Dynamic Routing.
2026-04-12 · 5 min read #ai-platform#cloudflare#ai-gateway#observability#cachingPreventing IT Outages & Lab Setup Guide + Global Table Tennis News Analysis
Key elements for preventing IT system outages, effective lab environment setup, and an analysis of global table tennis tournaments and notable players.
2026-04-11 · 14 min read #culture#devops#reliability#sre#table-tennisRAGAS Complete Guide: How to Quantitatively Evaluate Your RAG System
Learn to use RAGAS — the gold standard for RAG evaluation — to measure Faithfulness, Answer Relevancy, Context Precision, and Context Recall. Includes complete Python code, automated CI/CD evaluation pipelines, synthetic
2026-03-18 · 9 min read #ragas#rag-evaluation#llm-evaluation#ai-development#qualitySLI/SLO/Error Budget-Based Reliability Engineering: A Practical Guide
A comprehensive guide to reliability engineering with SLI/SLO/Error Budget. Covers SLI selection, SLO target setting, Error Budget policies, Burn Rate alerts, and Prometheus-based implementation to build a complete relia
2026-03-13 · 13 min read #observability#sli#slo#error-budget#sre