Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails Paper • 2609.09134 • Published 7 days ago • 10
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Paper • 2608.08722 • Published Aug 9 • 8
nvidia/NVIDIA-Nemotron-Labs-Teacher-Competition-Coding Text Generation • 561B • Updated Aug 14 • 1.56k • 6
view article Article Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL +6 aminediroHF, qgallouedec, kashif, lewtun, edbeeching, albertvillanova, lvwerra, sergiopaniego • May 27 • 50
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 15 days ago • 40
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published 19 days ago • 22
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published 25 days ago • 46
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published 29 days ago • 150
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 26 days ago • 275
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 27 days ago • 97
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published 28 days ago • 108