Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 21 days ago • 100
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 21 days ago • 186
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Paper • 2506.05209 • Published Jun 5, 2025 • 66
OlmoEarth Collection OlmoEarth pre-trained and fine-tuned foundation models for remote sensing • 18 items • Updated Jun 29 • 37
view article Article Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS nvidia • Aug 10 • 37
K2 Horizon Collection K2 Horizon models, datasets, and supporting resources • 22 items • Updated 12 days ago • 132
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Paper • 2607.25948 • Published Jul 28 • 16
Awesome feedback datasets Collection A curated list of datasets with human or AI feedback. Useful for training reward models or applying techniques like DPO. • 19 items • Updated Aug 3 • 71
view article Article Introducing EuroBERT: A High-Performance Multilingual Encoder Model EuroBERT • Mar 10, 2025 • 150
SWE-World: Building Software Engineering Agents in Docker-Free Environments Paper • 2602.03419 • Published Feb 3 • 42
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • Aug 10 • 42
Ornith-1.0 Collection Ornith-1.0 is a family of open-source LLMs specialized for agentic coding. • 8 items • Updated 3 days ago • 394
Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video Paper • 2510.03458 • Published Oct 3, 2025 • 4