Pretraining Datasets
The SPP pretraining corpus: reflections, the corpus selection manifest, safety scores, and verification data. Fully reproducible from public Dolma 3.
Viewer • Updated • 51.4MNote **The main artifact.** 51.4M documents paired with the first- and third-person constitution reflections the released models were trained on, plus the exact token offset each was inserted at. 155 GB.
dlab-spp/corpus-1T-manifest
Viewer • Updated • 1.06BNote Selection manifest for the 1T-token Dolma 3 subsample: which documents were chosen, in which order, and their safety scores. Rebuild the corpus from public data by document id — 16 GB instead of 2.6 TB.
dlab-spp/corpus-verification
Viewer • Updated • 1Note Megatron `.idx` document boundaries, per-document token counts and sha256s — verify an independently rebuilt corpus byte for byte without downloading the 2.17 TB token streams.
dlab-spp/safety-classifications
Viewer • Updated • 391M • 740Note SafeLM safety scores (0–5) and full class probabilities for 391M Dolma 3 documents. Publishing these removes the one stage of the pipeline that is not bit-reproducible.
dlab-spp/reflection-10m
Viewer • Updated • 10M • 694Note Earlier 10M-document reflection run, superseded by `reflection-50m`. Kept for the ablations that used it.
dlab-spp/reflection-sample-2k
Viewer • Updated • 2k • 67Note 2,000-row sample of the reflection data in the identical schema — for inspecting the format without downloading 155 GB.