Nouveau Recherche PDF, HTML, DOCX et bien d'autres formats.

Prépublication · 2026

Q-Learning with Scalar Adjoint Matching

Yonghoon Dong, Minsung Yoon et al. — Monde

Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the pol…

#cs.LG

Actions

Citation