Tag: #vlm
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 6 posts
Vision-Language Models (VLMs) 2026 Deep Dive — CLIP, LLaVA, InternVL3, Qwen2.5-VL, GPT-4o, Gemini 2.5, Claude 4.7, DINOv2, SAM 2, and Florence-2
Everything you need to know about Vision-Language Models in May 2026 in one place. CLIP family (SigLIP, EVA-CLIP), open VLMs (LLaVA-NeXT, InternVL3, Qwen2.5-VL, Pixtral, Molmo, Idefics3, MiniCPM-V), closed frontier (GPT-
2026-05-16 · 19 min read #vision-language-models#vlm#clip#llava#internvlComputer Vision Frameworks 2026 - OpenCV 4, MediaPipe, Detectron2, YOLO v11, MMDetection, SAM 2, Grounding DINO Deep Dive
The 2026 computer vision stack is no longer about "touching pixels". OpenCV 4.10 has made ONNX inference table stakes, MediaPipe Studio reduces mobile real-time pipelines to one line, YOLO v11 bundles NAS, segmentation,
2026-05-16 · 24 min read #computer-vision#opencv#mediapipe#detectron2#yoloThe 2026 Vision Model Development & Fine-Tuning Guide — CNN, ViT, DETR, SAM 2, VLMs and a Real Decision Tree
Vision model development in 2026 is no longer 'grab a ResNet and call it a day.' Between CNNs, ViTs, DETR variants, SAM 2, and VLMs like LLaVA, Qwen-VL, Gemini Vision, and Claude Vision, your choice for the same photo ca
2026-05-14 · 20 min read #computer-vision#vision-model#cnn#vit#detrThe Complete Guide to Multimodal LLMs: Vision, Document Understanding, OCR, Video, Audio, and the Specifics of Korean (2025)
The text-only era is over. In 2025, LLMs handle images, documents, video, and audio naturally. GPT-4o/Claude 3.5/Gemini/Qwen2-VL/Pixtral compared, Document AI and layout understanding, the modernization of OCR, video and
2026-04-15 · 14 min read #multimodal#vision-llm#document-ai#ocr#whisperLLM Multimodal Vision-Language Model Serving and Optimization Practical Guide
A practical guide to serving and optimizing multimodal vision-language models in production.
2026-03-05 · 24 min read #llm#multimodal#vlm#vllm#2026-03The Complete Autonomous Driving & Robotics Tech Stack: From C++, ROS2, CUDA, TensorRT to VLM/VLA, Simulation, and Beyond
A comprehensive guide to the core technology stack behind autonomous driving and robotics. Covering Modern C++, ROS/ROS2, CUDA parallel programming, TensorRT optimization, model compression (quantization/pruning), sensor
2026-03-01 · 22 min read #autonomous-driving#robotics#ros2#cuda#tensorrt