Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
guicybercode 's Collections
Papers I'm Reading
LLM Training from Scratch
Multilingual models
Smol LLMs for History

Papers I'm Reading

updated 13 days ago

Key papers on small LLMs, data-centric training, and multilingual NLP.

Upvote
-

  • The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

    Paper • 2406.17557 • Published Jun 25, 2024 • 105

  • SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

    Paper • 2502.02737 • Published Feb 4, 2025 • 261

  • The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

    Paper • 2506.05209 • Published Jun 5, 2025 • 64
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs