Tag: #llama-cpp
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 6 posts
llama.cpp Comes to the Browser — The Ceiling a WebGPU Backend Measured Across 16 Devices
LlamaWeb, the WebGPU backend for llama.cpp released by a UC Santa Cruz team in May 2026, left behind the widest dataset to date measuring browser LLM inference — 16 devices across 8 vendors. The results cut both ways. It
2026-07-16 · 16 min read #webgpu#llm-inference#llama-cpp#on-device-ai#browserLocal LLM Inference Optimization — From Quantization to Breaking the VRAM Ceiling
Privacy concerns, cost pressure, and big-tech fatigue are driving a local LLM revival. We map the entire landscape of local inference optimization: VRAM-first hardware thinking, GGUF and AWQ quantization, llama.cpp vs vL
2026-06-12 · 16 min read #llm#inference#quantization#llama-cpp#vllmLLM Serving & Local Inference in 2026 — vLLM / llama.cpp / MLX / Ollama / LM Studio / SGLang / TGI Deep Dive
A map of the 2026 LLM serving and inference landscape. Datacenter camp (vLLM, SGLang, TGI, Triton, TensorRT-LLM), local camp (llama.cpp, MLX, llamafile, Ollama, LM Studio, GPT4All), emerging camp (KTransformers, MLC LLM,
2026-05-16 · 25 min read #llm#model-serving#inference#vllm#llama-cppLocal AI & On-Device LLMs 2026 — Ollama · LM Studio · Jan · Msty · Open WebUI · GPT4All · AnythingLLM · Faraday Deep Dive
By May 2026, local AI is no longer a hobby. An M4 Max MacBook Pro runs Llama 4 Scout 109B MoE at 24 tokens per second. Desktop runtimes like Ollama, LM Studio, Jan, and Msty unify GUI and CLI, while Open WebUI, AnythingL
2026-05-16 · 23 min read #local-ai#on-device-llm#ollama#lm-studio#janEdge AI & TinyML 2026 — LiteRT / ExecuTorch / Edge Impulse / Jetson / Coral / Hailo / Sipeed K230 / llama.cpp / Phi-4 Deep-Dive Guide
A full-stack map of the 2026 Edge AI / TinyML ecosystem — the dual standard formed after TFLite Micro was rebranded as LiteRT and ExecuTorch reached GA, the TinyML cloud workflow created by Edge Impulse, the accelerator
2026-05-16 · 32 min read #edge-ai#tinyml#tflite-micro#litert#executorchAI Inference Engines 2026 - vLLM · SGLang · llama.cpp · TGI · TensorRT-LLM · MLX · mistral.rs · DeepSpeed-MII · Aphrodite Deep Dive
In 2026, LLM engineering is no longer about which model — it is about which inference engine. We dissect vLLM V1, SGLang 0.4, TensorRT-LLM, TGI 3.x, llama.cpp, MLX-LM, mistral.rs, DeepSpeed-MII, Aphrodite, CTranslate2, E
2026-05-16 · 20 min read #llm-inference#vllm#sglang#llama-cpp#tgi