The paper asks a simple question: when an AI sees a picture and some text but the instructions say 'only trust the picture,' how does it decide which one to follow?
This survey turns model understanding into a step-by-step repair toolkit called Locate, Steer, and Improve.
Large language models (LLMs) are good at many math problems but often mess up simple counting when the list gets long.