Tag: #ai-infrastructure
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 11 posts
Data Center Infrastructure Investing — The Picks and Shovels of the AI Buildout
Where does the real upside of the AI buildout live. This piece breaks down the data center value chain (real estate/REITs, cooling, power, networking, servers) through a picks-and-shovels lens, weighing both bull and bea
2026-06-18 · 28 min read #finance#data-center#ai-infrastructure#reit#coolingPower Grids, Copper, and Commodities — The Hidden Beneficiaries of the AI Power Buildout
The AI data center boom is lifting demand not only for chips but for the physical backbone behind them: power grids, copper, and transformers. This article examines the commodity cycle, supply constraints, and the inflat
2026-06-18 · 17 min read #commodities#copper#power-grid#ai-infrastructure#energyNvidia at 5 Trillion — How Far Along Is the AI Infrastructure Cycle?
Nvidia has crossed a 5 trillion dollar market capitalization for the first time ever. This article uses data to gauge where the AI infrastructure capex cycle stands today, weighing the bull and bear cases, supply-chain b
2026-06-18 · 8 min read #finance#nvidia#ai-infrastructure#capex#semiconductorsPower and Cooling in the AI Data Center — Infrastructure for the Gigawatt Era
As AI capex explodes and data centers grow to gigawatt scale, power and cooling have emerged as the dominant constraints. We map the big picture: surging rack power density, the limits of air cooling versus direct liquid
2026-06-16 · 21 min read #datacenter#ai-infrastructure#cooling#power#liquid-coolingVector Database Complete Guide 2025: Embeddings, Similarity Search, Pinecone/Weaviate/Qdrant/pgvector
Everything about Vector DBs! Vector embedding fundamentals, similarity search (cosine/euclidean/dot product), indexing algorithms (HNSW/IVF/PQ), Pinecone vs Weaviate vs Qdrant vs Milvus vs pgvector comparison, hybrid sea
2026-04-13 · 23 min read #vector-database#embedding#similarity-search#pinecone#weaviateLiteLLM Complete Guide 2025: Unify 100+ LLMs with a Single API Proxy Server
Everything about LiteLLM! 100+ LLM unified API, OpenAI-compatible proxy server, cost tracking/budget management, load balancing/fallback, model routing, virtual keys, rate limiting, Guardrails, production deployment (Doc
2026-03-25 · 18 min read #litellm#llm#api#proxy#openaiWEKA High-Performance Storage Complete Guide 2025: Parallel File System for AI/HPC Workloads
Everything about WEKA! Parallel file system architecture, NVMe tiering, GPU Direct Storage, AI/ML workload optimization, cloud integration (AWS/Azure/GCP), vs Ceph/Lustre/GPFS, data pipelines, performance benchmarks.
2026-03-25 · 22 min read #weka#wekafs#storage#parallel-filesystem#ai-infrastructureScale AI and the World of Data Labeling: Complete Guide to AI Training Data Industry and Careers
A comprehensive analysis of the AI training data industry led by Scale AI ($14B valuation). Covering data labeling principles, RLHF data pipelines, Scale AI vs Labelbox vs Snorkel comparison, data quality management, aut
2026-03-23 · 20 min read #scale-ai#data-labeling#annotation#rlhf#ai-training-dataGPU Hardware Complete Guide for AI: From Architecture to Selection Criteria
A comprehensive guide to GPU hardware for AI research and training. Covers NVIDIA GPU architectures (Hopper, Blackwell), Tensor Core, NVLink, HBM memory, A100/H100/H200/B200 comparisons, and cloud GPU options in detail.
2026-03-17 · 23 min read #gpu#hardware#nvidia#cuda#gpu-cudaAI System Design Complete Guide: From LLM Services to MLOps Architecture
A complete guide to designing production-grade AI systems. Learn real-world architectures for real-time inference systems, vector search infrastructure, LLM service architecture, data pipelines, and monitoring system des
2026-03-17 · 27 min read #system-design#ai-infrastructure#llm#mlops#architectureMastering Slurm: A Practical Guide to the HPC/AI Cluster Workload Manager
A comprehensive, hands-on guide to the Slurm workload manager. Covers architecture (slurmctld/slurmd/slurmdbd), core concepts (Partitions/QoS/Fairshare), essential commands (sbatch/srun/salloc), GPU scheduling (GRES/MIG/
2026-03-01 · 15 min read #slurm#hpc#gpu#distributed-training#cluster