How I Study AI - Learn AI Papers & Lectures the Easy Way

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning

Intermediate

Minwu Kim, Safal Shrestha et al.Jan 28arXiv

When training smart language models with RL that use right-or-wrong rewards, learning can stall on 'saturated' problems that the model almost always solves.

#failure-prefix conditioning#RLVR#GRPO

Papers1

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning