Blog
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3517 posts
#2026-03 765#english 592#culture 264#deep-dive 254#kubernetes 247#career 229#ai 216#llm 208#devops 193#2026-04 146#security 141#database 114#observability 113#communication 109#history 107#architecture 100#productivity 96#finance 88#economy 84#mindset 81#psychology 80#ai-papers 79#food 78#it 78#travel 78#deep-learning 77#japanese 77#networking 77#performance 72#business-travel 70#linux 70#gpu 69#ai-agent 66#cs-fundamentals 63#postgresql 60#rag 58#self-improvement 55#learning 53#mlops 53#ai-platform 51
Analyzing SOTA Multimodal LLMs — One Model to See, Hear, and Speak
How did a language model trained purely on text come to understand and generate images, audio, and video? This post walks through modality encoders and projectors, the unified token space, the any-to-any flow, native mul
2026-06-30 · 22 min read #multimodal-llm#any-to-any#vision-language#audio#architectureSOTA Speech Recognition and Synthesis — From Whisper to Codec Language Models
We survey recent trends in speech recognition (ASR) and synthesis (TTS). From HMM to CTC/attention, Whisper large-scale weak supervision, and from Tacotron to neural vocoders and codec language models, we trace the linea
2026-06-30 · 15 min read #ai-papers#speech-recognition#text-to-speech#whisper#neural-codecSOTA Autonomous Driving Perception — BEV, Occupancy, End-to-End
We survey recent trends in autonomous driving perception. From the perception-prediction-planning-control stack, BEV multi-camera fusion, occupancy networks 3D occupancy representation, vision-centric vs LiDAR fusion, to
2026-06-30 · 16 min read #ai-papers#autonomous-driving#bev#occupancy-network#end-to-endRobot Safety and Alignment — Trusting Powerful Robots
How can we trust increasingly powerful robots. We take a balanced look at physical safety, constrained reinforcement learning and safety layers, handling distribution shift, human-robot collaboration safety, verification
2026-06-29 · 17 min read #ai-papers#robotics#safety#alignment#reinforcement-learningRobots That Learn from Human Video — The Dream of Web-Scale Data
Can a robot learn from videos of people handling objects. We cover affordances and trajectories, the domain gap, representation learning and pre-training, one-shot imitation, web-video scale-up, and combining with robot
2026-06-29 · 19 min read #ai-papers#robotics#imitation-learning#representation-learning#videoRobot Foundation Models — One Policy for Many Jobs
We lay out the trend of robot foundation models that aim to do many jobs across many robots with a single policy. We cover the concept of a generalist policy, large-scale robot data such as Open X-Embodiment, cross-embod
2026-06-29 · 19 min read #ai-papers#robotics#foundation-model#generalist-policy#open-x-embodimentWorld Models — When Robots Imagine the Future
An explanation of the concept of world models (learning environment dynamics), covering model-based reinforcement learning, latent-space prediction, the robotic application of video prediction and generative models, plan
2026-06-29 · 17 min read #ai-papers#robotics#world-models#model-based-rl#planningSim-to-Real — Bringing What Was Learned in Simulation into Reality
An analysis of why robot policies learned in simulation collapse in reality (the reality gap), covering countermeasures such as domain randomization, domain adaptation, system identification, and digital twins, along wit
2026-06-29 · 17 min read #ai-papers#robotics#sim-to-real#domain-randomization#digital-twinHow Robots Learn — Imitation Learning and Reinforcement Learning
An overview of the four ways robots acquire skills, followed by a deep look at imitation learning (teleoperation, behavioral cloning, DAgger) and reinforcement learning (rewards, policies, exploration): their principles,
2026-06-29 · 16 min read #ai-papers#robotics#imitation-learning#reinforcement-learning#robot-learningThe Robots Eye — 3D Perception and SLAM
How robots see the world. We walk through RGB-D, LiDAR, and stereo sensors, the SLAM pipeline, point clouds and voxels, 6D pose estimation, deep-learning perception, and finally NeRF and 3D Gaussian Splatting, with code
2026-06-29 · 18 min read #ai-papers#robotics#slam#3d-perception#pointcloudHumanoid Whole-Body Control — Walking on Two Legs and Handling with Two Hands
From walking on two legs to handling objects with hands, we lay out the core ideas of bipedal locomotion and whole-body control. We cover ZMP and MPC, reinforcement-learning locomotion, balance and fall recovery, the int
2026-06-29 · 21 min read #ai-papers#robotics#humanoid#locomotion#whole-body-controlRegex From Zero to Confident: Character Classes, Quantifiers, Anchors, Groups, Lookarounds, and ReDoS
If regular expressions have always looked like line noise, let this post fix that. We build up the pieces one at a time — character classes, quantifiers, anchors, groups, alternation, lookarounds — then cover greedy vers
2026-06-29 · 10 min read #regex#programming#fundamentalsAsync Rust: async/await and the Tokio Runtime
Async in Rust has a different flavor from other languages. A Future is just a lazy state machine, and without an executor to poll it, nothing happens. What a Future really is, what .await means, the executor/runtime (Tok
2026-06-29 · 11 min read #rust#async#tokioA DNS Deep Dive — The Infrastructure Hidden Behind Names
DNS is not a simple name tag but a vast distributed infrastructure that holds up the internet. From the hierarchy and recursive/authoritative resolution, record types, TTL and caching, Anycast and GeoDNS, to DNSSEC and D
2026-06-29 · 22 min read #network#dns#anycast#cdn#infrastructureOpinionated Tools — The Philosophy of Code Formatters and Linters
Why code formatters end debates and preserve consistency. Comparing gofmt and gofumpt, prettier, black, and ruff, this piece lays out the philosophy of strictness versus flexibility, the difference from linters, CI enfor
2026-06-29 · 22 min read #devops#formatter#linter#developer-tools#toolingThe Rise of AI Code Review Tools — How Automated Review Changes Teams
AI code review tools are pouring out as open source and reshaping developer workflows. From Git diff analysis, defect detection, and convention enforcement to the division of labor with human reviewers, false-positive ma
2026-06-29 · 20 min read #devops#ai#code-review#ci-cd#developer-toolsDistributed Transactions: 2PC vs the Saga Pattern
Why the ACID that was easy inside one database falls apart across services, how two-phase commit works and where it breaks (coordinator, blocking, failure modes), the Saga pattern (choreography vs orchestration, compensa
2026-06-29 · 12 min read #distributed-systems#transactions#sagaTactile Sensing and Dexterous Manipulation — Robots That Feel with Their Fingertips
Robots have begun to feel the world with their fingertips. We lay out vision-based tactile sensors (the GelSight family) and electronic skin, the role of touch in sensing slip, force, and texture, visual-tactile fusion,
2026-06-29 · 21 min read #ai-papers#robotics#tactile-sensing#manipulation#gelsightJLPT N3 Reading and Listening — Strategies for a High Score
A focused attack on JLPT N3 reading and listening — the high-weight sections that trip up learners from Korean. Passage types and approaches, tracking keywords, referents, and connectors, time allocation, listening note-
2026-06-28 · 10 min read #japanese#jlpt#n3#reading#listeningJLPT N3 in 5 Days — An Ultra-Compressed Crash Course When Time Is Short
An emergency pass strategy for when you have only five days left before the exam. Give up on perfection and focus on clearing the 95-point pass line and avoiding a sectional fail. Includes a time-blocked 5-day plan, a mi
2026-06-28 · 28 min read #japanese#jlpt#n3#crash-course#study-guide