KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
IntermediateWuyang Zhou, Yuxuan Gu et al.Jan 29arXiv
Hyper-Connections (HC) make the usual single shortcut in neural networks wider by creating several parallel streams and letting the model mix them, but this can become unstable when stacked deep.
#Hyper-Connections#Manifold-Constrained Hyper-Connections#Doubly Stochastic Matrix