Papers2

#Vision Transformer (ViT)

Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations

This paper introduces NExT-Vid, a way to teach a video model by asking it to guess the next frame of a video while parts of the past are hidden.

#autoregressive video pretraining#masked next-frame prediction#context isolation

Not triaged yet

Stronger Normalization-Free Transformers

Intermediate

Mingzhi Chen, Taiming Lu et al.Dec 11arXiv

This paper shows that we can remove normalization layers from Transformers and still train them well by using a simple point‑by‑point function called Derf.

#Normalization‑free Transformers#LayerNorm replacement#Point‑wise activation

Not triaged yet