How I Study AI - Learn AI Papers & Lectures the Easy Way

Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning

Intermediate

Ming Chen, Sheng Tang et al.Dec 6arXiv

The paper shows that making a model write a number as a sequence of digits and then grading the whole number at the end works better than grading each digit separately.

#decoding-based regression#sequence-level reward#reinforcement learning

Papers1

Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning