DR-LoRA: Dynamic Rank LoRA for Mixture-of-Experts Adaptation
IntermediateGuanzhi Deng, Bo Li et al.Jan 8arXiv
Mixture-of-Experts (MoE) models use many small specialist networks and only activate a few per token, but classic LoRA fine-tuning gives every expert the same rank, wasting parameters on the wrong experts.
#DR-LoRA#Mixture-of-Experts#Low-Rank Adaptation