Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
IntermediateJi Zhao, Yufei Gu et al.Feb 5arXiv
Big idea: use a small, already-trained model to help a bigger model learn good habits early, so the big one trains faster and ends up smarter.
#Late-to-Early Training#LLM pretraining acceleration#representation alignment