Buckets:

182 GB
53,762 files
Updated 5 days ago
Name
Size
chunkmeta
chunks
films
meta
thumbmeta
thumbs
README.md1.55 kB
xet
skip_ids.parquet96.1 kB
xet
README.md

prelinger-films

Mirror of the explicitly public-domain subset of the Prelinger Archives (Internet Archive), prepared for batch video captioning with uv-scripts/video (Marlin-2B on HF Jobs).

Selection: collection:prelinger AND mediatype:movies AND licenseurl:*publicdomain* — 1,876 items, ~381 hours. Every item carries an explicit Creative Commons public domain mark in its IA metadata; see each film's sidecar for the per-item rights basis. Context: the archive states ~65% of its holdings are US public domain, and materials hosted on IA carry a standing public-domain dedication supporting commercial and non-commercial reuse. PD status is a US determination.

Layout

  • films/{identifier}.mp4 — best available MP4 derivative per item (512kb preferred)
  • meta/{identifier}.json — per-film sidecar: identifier, licenseurl, title, date, runtime, source URL, fetch timestamp. Sidecar-exists == film fully copied.
  • chunks/{identifier}/c%04d.mp4 — ~60s stream-copy segments (keyframe-snapped)
  • chunkmeta/{identifier}.json — actual segment boundary times (offsets for timestamp remapping are these recorded values, not assumed 60s multiples)
  • captions/ — saturate output: resumable parquet, one row per chunk (scene, caption, events JSON with global-time <start – end> events)

Built 2026-07-30 with HF Jobs; pipeline scripts archived in hf://buckets/davanstrien/prelinger-sample/scripts/.

Total size
182 GB
Files
53,762
Last updated
Jul 30
Pre-warmed CDN
US EU US EU

Contributors