研究视觉模型在旋转和噪声下关系推理的错误,揭示其脆弱性。
When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise

- 测试旋转与噪声对视觉语言模型关系推理的影响
- 轻微扰动即导致关系误判率显著上升
- 现有增强方法改善有限,需更几何敏感的模型
视觉语言模型(VLMs)虽具备强大多模态能力,但仍易产生关系幻觉,需准确推理物体间交互。本文研究旋转与噪声等视觉扰动的影响,发现即使轻微失真也会显著降低各类模型与数据集的关系推理性能。进一步评估了基于提示的增强与预处理策略(方向校正与去噪),结果表明这些方法仅带来部分改进,无法完全解决幻觉问题。研究揭示感知鲁棒性与关系理解之间存在差距,强调需要更稳健、具备几何意识的视觉语言模型。
原文摘要 · Abstract (English)
Vision-language models (VLMs) achieve strong multimodal performance but remain prone to relation hallucination, which requires accurate reasoning over inter-object interactions. We study the impact of visual perturbations, specifically rotation and noise, and show that even mild distortions significantly degrade relational reasoning across models and datasets. We further evaluate prompt-based augmentation and preprocessing strategies (orientation correction and denoising), finding that while they offer partial improvements, they do not fully resolve hallucinations. Our results reveal a gap between perceptual robustness and relational understanding, highlighting the need for more robust, geometry-aware VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。