Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models Paper • 2608.16647 • Published Aug 17 • 15
RLVR Linearity Collection RL training and evaluation datasets, and checkpoints in 'Linear Dynamics in the RLVR Training of Large Language Models' • 3 items • Updated May 22
RLVR Linearity Collection RL training and evaluation datasets, and checkpoints in 'Linear Dynamics in the RLVR Training of Large Language Models' • 3 items • Updated May 22