Tag: #ray
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 4 posts
Label Locality Scheduling in Ray 2.56 — Placement Groups Start Seeing NVLink Racks, Not Nodes
Ray 2.56.0, released on June 29, 2026, adds a domain-level scheduling layer to placement groups as an alpha feature. Until now every placement strategy — PACK, STRICTPACK and friends — operated strictly at node granulari
2026-07-17 · 11 min read #ray#gpu#scheduling#distributed-systemsMLOps Platforms 2026 Deep Dive — MLflow, Kubeflow, W&B, Vertex AI, SageMaker, Databricks, BentoML, Ray, Modal, Hugging Face
A side-by-side look at 30+ MLOps platforms in May 2026. MLflow 3, Kubeflow 1.10, Weights & Biases, Comet, Neptune.ai, ClearML, Vertex AI, SageMaker, Azure ML, Databricks ML + Mosaic AI, Hugging Face Inference Endpoints,
2026-05-16 · 17 min read #mlops#mlflow#kubeflow#weights-and-biases#vertex-aiDistributed Training & GPU Infrastructure 2026 Deep-Dive — DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, Blackwell GB200, MI325X, TPU v5p
A comparison of DeepSpeed, FSDP2, Megatron-LM, Ray Train, JAX, TorchTitan, and Composer — plus NVIDIA Blackwell GB200 NVL72, AMD MI325X, Intel Gaudi 3, AWS Trainium 2, and Google TPU v5p/v6e Trillium. 3D parallelism, ZeR
2026-05-16 · 14 min read #distributed-training#deepspeed#fsdp#megatron-lm#rayMLOps Complete Guide — Model Serving, Feature Store, Drift, A/B Testing, GPU Economics (Season 2 Ep 7, 2025)
Training a model and running it in production are completely different games. Serving (TorchServe, Triton, vLLM, TGI), Feature Stores (Feast, Tecton), training infra (Ray, Determined), experiment tracking (MLflow, W&B),
2026-04-15 · 12 min read #mlops#model-serving#feature-store#drift-detection#ab-testing