Prépublication · 2026
A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
Randy Ardywibowo, Arnav Dalal et al. — Monde
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting…
#cs.LG