Papers5

#on-policy training

T3D: Few-Step Diffusion Language Models via Trajectory Self-Distillation with Direct Discriminative Optimization

Tunyu Zhang, Xinxi Zhang et al.Feb 12arXiv

This paper shows how to make diffusion language models write high‑quality text in just a few steps, which makes them much faster.

#diffusion language models#few-step decoding#trajectory self-distillation

Not triaged yet

On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models

Intermediate

Shumin Wang, Yuexiang Xie et al.Feb 3arXiv

The paper builds a simple, math-light rule to predict whether training makes a language model more open-minded (higher entropy) or more sure of itself (lower entropy).

#reinforcement fine-tuning#entropy dynamics#GRPO

Not triaged yet

EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience

Intermediate

Taofeng Xue, Chong Peng et al.Jan 22arXiv

Before this work, computer-using AIs mostly copied old examples and struggled with long step-by-step tasks on real computers.

#computer use agent#verifiable synthesis#validator

Not triaged yet

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

Beginner

Zhiwei Zhang, Fei Zhao et al.Jan 22arXiv

Small AI models often stumble when a tool call fails and then get stuck repeating bad calls instead of fixing the mistake.

#FISSION-GRPO#error recovery#tool use

Not triaged yet

Diversity or Precision? A Deep Dive into Next Token Prediction

Intermediate

Haoyuan Wu, Hai Wang et al.Dec 28arXiv

The paper shows that teaching a language model with a special “reward-shaped” next-token objective can make later reinforcement learning (RL) work much better.

#next-token prediction#cross-entropy as policy gradient#reward shaping

Not triaged yet