arXiv:2607.22600cs.AI2026-07

测试视觉语言模型对误导性图表的脆弱性并提出缓解方法。

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

论文配图:Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
图 1 · 摘自论文原文
  • 构建首个成对对比的误导图表基准,涵盖8类欺骗设计。
  • 32,000次测试显示顶尖模型仍易受误导,平均错误率超40%。
  • 提出推理时多智能体框架,利用图表元数据提升判断可靠性。

信息可视化广泛用于传达模式、趋势和异常值,但截断或反转坐标轴、扭曲比例、不当编码和误导性颜色映射等欺骗性设计,可在不改变原始数据的情况下系统性改变解读。随着视觉语言模型(VLMs)在图表理解与分析推理中的应用日益增多,评估其对这类误导性可视化的鲁棒性已成为可信数据分析的关键。本文提出VisDeception,首个受控成对基准,用于评估VLMs对误导图表设计的鲁棒性。该基准包含1,600对图表,涵盖8大类误导可视化手法,每张误导图表均来自同一组基础数据生成的忠实对照图。为分离欺骗引发的推理错误与基础图表理解误差,引入了“欺骗得分”这一配对评估指标,量化误导可视化使模型回答偏离真实数据解释的程度。在10个先进VLMs上进行的32,000次测试表明,即使是最先进的模型也对误导性视觉操纵高度敏感。为此,进一步提出一种推理时的多智能体缓解框架,通过在生成答案前提取结构化图表元数据来锚定推理过程,从而在无需用户显式指令的情况下降低误导视觉线索的影响。研究揭示当前图表理解系统存在重要可靠性缺陷,并确立基于基准的评估、欺骗感知指标与结构化推理作为发展更可信视觉分析VLMs的可行方向。

原文摘要 · Abstract (English)

Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data. As Vision-Language Models (VLMs) are increasingly used for chart understanding and analytical reasoning, assessing their robustness to such deceptive visualizations has become critical for trustworthy data analysis. We introduce VisDeception, the first controlled paired benchmark for evaluating the robustness of VLMs to misleading chart designs. The benchmark contains 1,600 paired faithful and misleading charts spanning eight major categories of deceptive visualization tactics, where each misleading chart is paired with a faithful counterpart generated from the same underlying data. To isolate deception-induced reasoning errors from baseline chart-understanding errors, we introduce the Deception Score, a paired evaluation metric that quantifies how misleading visualizations shift model responses away from the faithful interpretation of the data. Across 32,000 responses from 10 state-of-the-art VLMs, we find that even advanced models remain highly vulnerable to deceptive visual manipulations. To improve robustness, we further propose an inference-time multi-agent mitigation framework that grounds reasoning in structured chart metadata extracted from the visualization before answer generation, enabling models to reduce the influence of deceptive visual cues without requiring explicit user instructions. Together, our findings reveal important reliability gaps in current chart-understanding systems and establish benchmark-driven evaluation, deception-aware metrics, and structured reasoning as promising directions for developing more trustworthy VLMs for visual analytics.

视觉语言模型图表理解误导可视化鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。