Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| chunkmeta | 1,864 items | ||
| chunks | 23,148 items | ||
| films | 1,868 items | ||
| meta | 1,868 items | ||
| thumbmeta | 1,864 items | ||
| thumbs | 23,148 items | ||
| README.md | 1.55 kB xet | 47db8ff1 | |
| skip_ids.parquet | 96.1 kB xet | b69fb4a5 |
prelinger-films
Mirror of the explicitly public-domain subset of the Prelinger Archives (Internet Archive), prepared for batch video captioning with uv-scripts/video (Marlin-2B on HF Jobs).
Selection: collection:prelinger AND mediatype:movies AND licenseurl:*publicdomain*
— 1,876 items, ~381 hours. Every item carries an explicit Creative Commons public
domain mark in its IA metadata; see each film's sidecar for the per-item rights basis.
Context: the archive states ~65% of its holdings are US public domain, and materials
hosted on IA carry a standing public-domain dedication supporting commercial and
non-commercial reuse. PD status is a US determination.
Layout
films/{identifier}.mp4— best available MP4 derivative per item (512kb preferred)meta/{identifier}.json— per-film sidecar:identifier,licenseurl,title,date,runtime, source URL, fetch timestamp. Sidecar-exists == film fully copied.chunks/{identifier}/c%04d.mp4— ~60s stream-copy segments (keyframe-snapped)chunkmeta/{identifier}.json— actual segment boundary times (offsets for timestamp remapping are these recorded values, not assumed 60s multiples)captions/— saturate output: resumable parquet, one row per chunk (scene,caption,eventsJSON with global-time<start – end>events)
Built 2026-07-30 with HF Jobs; pipeline scripts archived in
hf://buckets/davanstrien/prelinger-sample/scripts/.
- Total size
- 182 GB
- Files
- 53,762
- Last updated
- Jul 30
- Pre-warmed CDN
- US EU US EU