Tag: #video-generation
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 7 posts
Video Generation and Understanding Technical Reports: What to Read, and Why Constraints Beat Demos
Ten video generation and understanding technical reports, each verified by opening the arXiv abstract page directly. CogVideoX, Movie Gen, HunyuanVideo, LTX-Video, Wan, Seedance 1.0 and 2.0, plus Qwen2.5-VL, VideoLLaMA 3
2026-08-12 · 7 min read #ai-papers#paper-review#technical-report#video-generation#diffusion-transformerMaking Video from a Single Image — Kling·Veo·Sora vs Wan·HunyuanVideo, What to Pick and When
When you are choosing a model to turn a single image plus a prompt into video, what actually decides it is not the polish of the demo reel but three things: price per second, input-image constraints, and licensing. This
2026-07-17 · 24 min read #ai#video-generation#image-to-video#open-weights#licensingSOTA Video Generation Models Explained — Spatiotemporal Diffusion Transformers
A look at the two fundamental challenges of video generation, temporal consistency and compute cost, through the lens of spatiotemporal latent patches and diffusion transformers. We analyze the concepts Sora introduced,
2026-06-30 · 8 min read #ai-papers#video-generation#diffusion-transformer#spatiotemporal#text-to-videoAI Video Generation 2026 — Sora 2 / Veo 3 / Kling 2 / Hailuo / Runway Gen-4 / Luma Ray 2 / HunyuanVideo Deep Dive
Two years after Sora 1 stunned the film industry in Feb 2024, the AI video generation landscape of May 2026 has hardened around three camps — closed-source SOTA (Sora 2, Veo 3, Kling 2, Hailuo), industry-standard video t
2026-05-15 · 27 min read #ai-video#video-generation#sora#veo#kling2025 AI Research Trends: Top HuggingFace Papers and 10 Defining Research Directions
A developer-focused review of HuggingFace trending papers and the 10 defining AI research trends of 2025. DeepSeek-R1 pure RL reasoning, Nemotron-Cascade 30B/3B MoE, GRPO, PagedAttention, million-token context limitation
2026-03-21 · 15 min read #ai-research#papers#huggingface#reasoning#moeWan Text-to-Video/Image-to-Video and Z Image Turbo Complete Analysis: Architecture and Applications of Next-Gen Video/Image Generation Models
A complete analysis of Wan video generation models and Z Image Turbo covering architecture, performance, and practical applications.
2026-03-01 · 36 min read #wan#text-to-video#image-to-video#z-image-turbo#video-generationHunyuanVideo and LTX-2 Complete Analysis: Architecture, Performance, and Practical Guide to Open-Source Video Generation Models
A deep analysis of Tencent HunyuanVideo (13B) and Lightricks LTX-2 (19B) architectures, training methodologies, and performance benchmarks, along with a comprehensive comparison of the open-source video generation ecosys
2026-03-01 · 30 min read #hunyuan-video#ltx-video#ltx2#text-to-video#image-to-video