How I Study AI - Learn AI Papers & Lectures the Easy Way

Tool Verification for Test-Time Reinforcement Learning

Ruotong Liao, Nikolai Röhrich et al.Mar 2arXiv

The paper fixes a big flaw in test-time reinforcement learning (TTRL): when many wrong answers agree, the model rewards the mistake and gets stuck.

#test-time reinforcement learning#verification-weighted voting#tool verification

Not triaged yet

Self-Improving Pretraining: using post-trained models to pretrain better models

Intermediate

Ellen Xiaoqing Tan, Shehzaad Dhuliawala et al.Jan 29arXiv

This paper teaches language models to be safer, more factual, and higher quality during pretraining, not just after, by using reinforcement learning with a stronger model as a helper.

#self-improving pretraining#reinforcement learning#online DPO

Not triaged yet

Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models

Intermediate

Youwei Liu, Jian Wang et al.Jan 13arXiv

Agents often act like tourists without a map: they react to what they see now and miss long-term consequences.

#Imagine-then-Plan#world models#adaptive lookahead

Not triaged yet

Papers3

Tool Verification for Test-Time Reinforcement Learning

Self-Improving Pretraining: using post-trained models to pretrain better models

Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models