ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
BeginnerTao Liu, Taiqiang Wu et al.Jan 14arXiv
Traditional supervised fine-tuning (SFT) makes a model copy one answer too exactly, which can cause overfitting to the exact wording instead of the real idea.
#ProFit#Supervised Fine-Tuning#Token Probability