arXiv:2509.18425cs.CV2025-09被引 6

测试视觉语言模型在有缺陷图表上的表现,发现其易产生幻觉且自信错误。

Losing the Plot: How VLM responses degrade on imperfect charts

  • 构建包含噪声与遮挡的图表测试集,用反向不一致提示检测模型自相矛盾
  • GPT-4o、Claude Sonnet 4等模型在图表退化时准确率显著下降,幻觉率升至40%以上
  • 适合关注图表理解可靠性、AI可解释性的研究者与开发者

视觉语言模型(VLMs)在图表理解任务上表现优异,但现有基准假设图表清晰且问题基于事实。现实中的图表常含失真或遮挡,需超越简单匹配的推理能力。我们评估了ChatGPT 4o、Claude Sonnet 4和Gemini 2.5 Pro,发现其在图表退化条件下性能急剧下降,幻觉现象(如数值虚构、趋势误判、实体混淆)频率显著上升。模型在退化环境中仍高度自信,生成看似合理但无依据的解释。为填补此空白,我们提出CHART NOISe数据集,融合图表失真、遮挡及受韩国CSAT英语考试启发的多选题设计;核心创新在于提示反向不一致性——要求模型对同一陈述先确认后否认,暴露其内在矛盾。贡献包括:(1) 对前沿VLMs进行评测,揭示图表推理中的系统性漏洞;(2) 发布首个整合失真、遮挡与反向不一致的CHART NOISe数据集;(3) 提出质量过滤与遮挡检测等基线缓解策略。这些工作共同建立了更严格的图表理解鲁棒性评估体系。

原文摘要 · Abstract (English)

Vision language models (VLMs) show strong results on chart understanding, yet existing benchmarks assume clean figures and fact based queries. Real world charts often contain distortions and demand reasoning beyond simple matching. We evaluate ChatGPT 4o, Claude Sonnet 4, and Gemini 2.5 Pro, finding sharp performance drops under corruption or occlusion, with hallucinations such as value fabrication, trend misinterpretation, and entity confusion becoming more frequent. Models remain overconfident in degraded settings, generating plausible but unsupported explanations. To address this gap, we introduce CHART NOISe(Chart Hallucinations, Answers, and Reasoning Testing on Noisy and Occluded Input Selections), a dataset combining chart corruptions, occlusions, and exam style multiple choice questions inspired by Korea's CSAT English section. A key innovation is prompt reverse inconsistency, where models contradict themselves when asked to confirm versus deny the same statement. Our contributions are threefold: (1) benchmarking state of the art VLMs, exposing systematic vulnerabilities in chart reasoning; (2) releasing CHART NOISe, the first dataset unifying corruption, occlusion, and reverse inconsistency; and (3) proposing baseline mitigation strategies such as quality filtering and occlusion detection. Together, these efforts establish a rigorous testbed for advancing robustness and reliability in chart understanding.

图表理解幻觉检测鲁棒性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。