Tag: #common-crawl
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
Open Source AI Training Datasets in 2026 — Common Crawl / FineWeb (HF) / RedPajama-V2 / Dolma / SlimPajama / The Stack v2 / LAION / COYO-700M (Kakao) Deep Dive
A full atlas of open source AI training datasets in 2026. The foundation Common Crawl and its refined descendants RefinedWeb / RedPajama-V2 / FineWeb / FineWeb-Edu / Dolma / SlimPajama, the academic stacks The Pile / S2O
2026-05-16 · 18 min read #ai-datasets#training-data#common-crawl#refinedweb#redpajamaInternet Archive & Digital Preservation & Web Archiving 2026 — Wayback Machine / archive.today / Conifer / Browsertrix / WARC / Perma.cc / NDL WARP / OASIS Deep Dive
A May 2026 snapshot of the digital preservation and web archiving ecosystem. Covers Brewster Kahles Internet Archive and its 935B-page Wayback Machine, the anonymous archive.today, the Webrecorder family (Conifer, Browse
2026-05-16 · 26 min read #internet-archive#wayback-machine#brewster-kahle#archive-today#archive-ph