Tag: #diffusion-transformer
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 7 posts
Video Generation and Understanding Technical Reports: What to Read, and Why Constraints Beat Demos
Ten video generation and understanding technical reports, each verified by opening the arXiv abstract page directly. CogVideoX, Movie Gen, HunyuanVideo, LTX-Video, Wan, Seedance 1.0 and 2.0, plus Qwen2.5-VL, VideoLLaMA 3
2026-08-12 · 7 min read #ai-papers#paper-review#technical-report#video-generation#diffusion-transformerImage Generation and Understanding Technical Reports: What to Read, and Why Design Beats Sample Images
Eleven image generation and understanding technical reports, each verified by opening the arXiv abstract page directly. Rectified flow transformers, VAR, Emu3, SANA, Janus-Pro, FLUX.1 Kontext, the Qwen-Image line, Seedre
2026-08-12 · 8 min read #ai-papers#paper-review#technical-report#image-generation#diffusion-transformerSOTA Video Generation Models Explained — Spatiotemporal Diffusion Transformers
A look at the two fundamental challenges of video generation, temporal consistency and compute cost, through the lens of spatiotemporal latent patches and diffusion transformers. We analyze the concepts Sora introduced,
2026-06-30 · 8 min read #ai-papers#video-generation#diffusion-transformer#spatiotemporal#text-to-videoAnalyzing SOTA Image Generation Models — From Diffusion to FLUX
A lineage-centered overview of the frontier of text-to-image generation, from diffusion model fundamentals through latent diffusion, DiT, rectified flow, and the FLUX family. We analyze the shared structure and differenc
2026-06-30 · 10 min read #ai-papers#diffusion-models#text-to-image#latent-diffusion#rectified-flowFoundation Model Architectures 2026 — Beyond the Transformer / Mamba 2 / Hyena / RWKV / RetNet / Griffin / Jamba / xLSTM / TTT / DiT / MoE / Flash Attention 3 Deep Dive
In 2026 the foundation-model world is no longer Transformer-only. Vaswani 2017 "Attention is All You Need" remains the standard, but next to it stand state-space models (Mamba, Mamba 2), the linear-RNN renaissance (RWKV,
2026-05-16 · 22 min read #foundation-models#transformer#attention-is-all-you-need#vaswani#mambaDiffusion Transformer (DiT) Architecture Analysis: The Shift from U-Net to Transformer
An analysis of the Scalable Diffusion Models with Transformers (DiT) paper. We cover the motivations behind transitioning from U-Net backbones to Transformers, adaLN-Zero conditioning, scaling laws, and the downstream im
2026-03-03 · 10 min read #ai-papers#diffusion-transformer#dit#generative-ai#image-generationHunyuanVideo and LTX-2 Complete Analysis: Architecture, Performance, and Practical Guide to Open-Source Video Generation Models
A deep analysis of Tencent HunyuanVideo (13B) and Lightricks LTX-2 (19B) architectures, training methodologies, and performance benchmarks, along with a comprehensive comparison of the open-source video generation ecosys
2026-03-01 · 30 min read #hunyuan-video#ltx-video#ltx2#text-to-video#image-to-video