When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 6 days ago • 42
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 7 days ago • 61
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 15 days ago • 81
Temporal Preference Optimization for Unsupervised Retrieval Paper • 2606.17664 • Published Jun 16 • 1
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation Paper • 2603.11137 • Published Mar 11 • 1
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 6 days ago • 42