Learning to Discover at Test Time
IntermediateMert Yuksekgonul, Daniel Koceja et al.Jan 22arXiv
This paper shows how to keep training a language model while it is solving one hard, real problem, so it can discover a single, truly great answer instead of many average ones.
#test-time training#reinforcement learning#entropic objective