Large Vision-Language Models (LVLMs) are great with one picture but get confused when you give them several, often mixing details from different images.
Large Vision-Language Models (LVLMs) look great on single images but often stumble when they must reason across multiple images.