Multimodal AI models can mix up what they see and what they hear, making things up across senses; this is called cross-modal hallucination.
This paper teaches video-language models to first find when the proof happens in a video and then answer with that proof, instead of mixing both steps together.