This paper introduces Knot Forcing, a way to make talking-head videos that look great while being generated live, frame by frame.
InfiniteVL is a vision-language model that mixes two ideas: local focus with Sliding Window Attention and long-term memory with a linear module called Gated DeltaNet.