Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
IntermediateYingfa Chen, Zhen Leng Thai et al.Jan 29arXiv
This paper shows how to turn a big Transformer model into a faster hybrid model that mixes attention and RNN layers using far less training data (about 2.3B tokens).
#hybrid attention#RNN attention hybrid#linear attention