When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

📰 ArXiv cs.AI

arXiv:2507.13868v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual input can lead to hallucinations and unreliable predictions. In this work, we investigate the mechanisms that VLMs use to resolve cross-modal conflicts by introducing WHOOPS-AHA!, a dataset of multimodal counterfactual queries that deliberately contradict intern

Published 21 Apr 2026
Read full paper → ← Back to Reads