This paper shows how to turn a big Transformer model into a faster hybrid model that mixes attention and RNN layers using far less training data (about 2.3B tokens).
The paper introduces Canon layers, tiny add-ons that let nearby words share information directly, like passing notes along a row of desks.