β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 5 days ago • 21