Papers4

#GPQA-Diamond

CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning

CHIMERA is a small (about 9,000 examples) but very carefully built synthetic dataset that teaches AI to solve hard problems step by step.

#CHIMERA dataset#synthetic data generation#chain-of-thought

Not triaged yet

ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection

Beginner

Tao Liu, Taiqiang Wu et al.Jan 14arXiv

Traditional supervised fine-tuning (SFT) makes a model copy one answer too exactly, which can cause overfitting to the exact wording instead of the real idea.

#ProFit#Supervised Fine-Tuning#Token Probability

Not triaged yet

Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling

Beginner

Falcon LLM Team, Iheb Chaabane et al.Jan 5arXiv

Falcon-H1R is a small (7B) AI model that thinks really well without needing giant computers.

#Falcon-H1R#Hybrid Transformer-Mamba#Chain-of-Thought

Not triaged yet

Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory

Beginner

Mirac Suzgun, Mert Yuksekgonul et al.Apr 10arXiv

The paper introduces Dynamic Cheatsheet (DC), a simple way for language models to keep a tiny, smart notebook of useful tricks while they are being used.

#Dynamic Cheatsheet#test-time learning#memory curation

Not triaged yet