This paper introduces NExT-Vid, a way to teach a video model by asking it to guess the next frame of a video while parts of the past are hidden.
This paper shows that we can remove normalization layers from Transformers and still train them well by using a simple point‑by‑point function called Derf.