Factorized Learning for Temporally Grounded Video-Language Models
IntermediateWenzheng Zeng, Difei Gao et al.Dec 30arXiv
This paper teaches video-language models to first find when the proof happens in a video and then answer with that proof, instead of mixing both steps together.
#temporal grounding#video-language models#evidence tokens