Nouveau Recherche PDF, HTML, DOCX et bien d'autres formats.

Prépublication · 2026

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Randy Ardywibowo, Arnav Dalal et al. — Monde

Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting…

#cs.LG

Actions

Citation