Enhancing Multi-Image Understanding through Delimiter Token Scaling
IntermediateMinyoung Lee, Yeji Park et al.Feb 2arXiv
Large Vision-Language Models (LVLMs) are great with one picture but get confused when you give them several, often mixing details from different images.
#Large Vision-Language Models#Multi-image understanding#Delimiter tokens