Zh-Pythia Collection A series of Chinese language models trained on 3B tokens • 5 items • Updated Nov 13, 2024 • 1
Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures Paper • 2510.24081 • Published Oct 28, 2025 • 24
VISTA 130M Collection VISTA 130M models from "No Single Tokenizer Feature Reliably Predicts Downstream Language Model Performance" (EMNLP 2026, Findings) • 84 items • Updated 15 days ago • 3
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Paper • 2506.05209 • Published Jun 5, 2025 • 65
view article Article Releasing the largest multilingual open pretraining dataset Pclanglais • Nov 13, 2024 • 110