Tag: #training-data
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
Open Source AI Training Datasets in 2026 — Common Crawl / FineWeb (HF) / RedPajama-V2 / Dolma / SlimPajama / The Stack v2 / LAION / COYO-700M (Kakao) Deep Dive
A full atlas of open source AI training datasets in 2026. The foundation Common Crawl and its refined descendants RefinedWeb / RedPajama-V2 / FineWeb / FineWeb-Edu / Dolma / SlimPajama, the academic stacks The Pile / S2O
2026-05-16 · 18 min read #ai-datasets#training-data#common-crawl#refinedweb#redpajamaComplete Guide to Korean LLM Training Data: Hugging Face Datasets, Preprocessing, and Quality Control
Everything about LLM training data! Hugging Face datasets (types/loading/conversion), Korean data collection (crawling/synthetic/translation), preprocessing (tokenization/cleaning/dedup), Instruction Tuning formats (Alpa
2026-03-25 · 23 min read #llm#training-data#huggingface#dataset#korean-nlp