Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
IntermediateMing Chen, Sheng Tang et al.Dec 6arXiv
The paper shows that making a model write a number as a sequence of digits and then grading the whole number at the end works better than grading each digit separately.
#decoding-based regression#sequence-level reward#reinforcement learning