See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
IntermediateLe Thien Phuc Nguyen, Zhuoran Yu et al.Dec 1arXiv
This paper introduces AV-SpeakerBench, a new test that checks if AI can truly see, hear, and understand who is speaking, what they say, and when they say it in real videos.
#audiovisual reasoning#speaker attribution#temporal grounding