Tag: #sre
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 23 posts
Complete Guide to English Incident Communication: From Outage Reports to Postmortem Writing
An English incident communication guide for IT engineers. Covers initial incident reporting, escalation, status page updates, and postmortem writing with practical expressions and templates.
2026-03-08 · 26 min read #english#incident-communication#postmortem#sre#technical-writingJapanese for SRE: On-Call Alert Response, Incident Escalation, and Post-Mortem Communication
Japanese OUTPUT training for SRE and DevOps engineers covering on-call alert acknowledgment, incident escalation communication, status update phrases, post-mortem discussion, and keigo usage in technical emergency contex
2026-03-07 · 30 min read #japanese#sre#on-call#incident-response#communicationThe Complete Guide to Redundancy — The Secret Behind 99.999% Availability
From dual redundancy to triple redundancy, Active-Standby, Active-Active, N+1, split brain, and quorum. We explain why the difference between 99.9% and 99.999% translates to 8 hours vs. 5 minutes of annual downtime, with
2026-03-02 · 8 min read #architecture#high-availability#redundancy#failover#load-balancing