HERMES is a training-free way to make video-language models understand live, streaming video quickly and accurately.
InfiniteVL is a vision-language model that mixes two ideas: local focus with Sliding Window Attention and long-term memory with a linear module called Gated DeltaNet.