Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
IntermediateKirill Pavlenko, Alexander Golubev et al.Feb 10arXiv
The paper fixes a common mistake in training language models for multi-part tasks: giving the same reward signal to every token, even when different text parts aim at different goals.
#Blockwise Advantage Estimation#Outcome-Conditioned Baseline#Group Relative Policy Optimization