More Images, More Problems? A Controlled Analysis of VLM Failure Modes
IntermediateAnurag Das, Adrian Bulat et al.Jan 12arXiv
Large Vision-Language Models (LVLMs) look great on single images but often stumble when they must reason across multiple images.
#Large Vision-Language Models#Multi-image reasoning#Cross-image aggregation