CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
IntermediateZhiyuan Yao, Yi-Kai Zhang et al.Feb 3arXiv
Large language models learn better when we spend more practice time on the right questions at the right moments.
#Reinforcement Learning#RLVR#GRPO