How I Study AI - Learn AI Papers & Lectures the Easy Way

RIVER: A Real-Time Interaction Benchmark for Video LLMs

Intermediate

Yansong Shi, Qingsong Zhao et al.Mar 4arXiv

RIVER Bench is a new test that checks how well AI can watch a video stream and talk with you in real time.

#RIVER Bench#online video understanding#multimodal large language models

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Intermediate

Christopher Clark, Jieyu Zhang et al.Jan 15arXiv

Molmo2 is a family of vision-language models that can watch videos, understand them, and point to or track things over time using fully open weights, data, and code.

#vision-language model#video grounding#pointing and tracking

A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos

Intermediate

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan et al.Dec 18arXiv

This paper builds a new test, LongShOTBench, to check if AI can truly understand long videos by using sight, speech, and sounds together.

#long-form video understanding#multimodal reasoning#audio-visual-speech alignment

Papers3

RIVER: A Real-Time Interaction Benchmark for Video LLMs

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos