How I Study AI - Learn AI Papers & Lectures the Easy Way

ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection

Beginner

Tao Liu, Taiqiang Wu et al.Jan 14arXiv

Traditional supervised fine-tuning (SFT) makes a model copy one answer too exactly, which can cause overfitting to the exact wording instead of the real idea.

#ProFit#Supervised Fine-Tuning#Token Probability

Papers1

ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection