arXiv:2503.18172cs.CLcs.AI2025-03EMNLP被引 20

评测大模型识破误导性图表的能力,发现其存在明显漏洞。

Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering

  • 构建3026个案例的误导图表数据集,覆盖21类造假手法。
  • 24个主流大模型平均准确率不足50%,在柱状图造假上表现最差。
  • 提出区域感知推理框架,显著提升识别能力,适合可信AI研究者使用。

误导性可视化通过操纵图表呈现方式来支持特定观点,扭曲认知并导致错误结论。尽管数十年研究持续进行,此类问题仍普遍存在,威胁公众理解力,并对参与数据传播的AI系统构成安全风险。虽然近期多模态大语言模型(MLLM)展现出较强的图表理解能力,但其识别与解析误导性图表的能力尚未被探索。本文提出Misleading ChartQA基准,一个大规模多模态数据集,用于评估MLLM在误导图表推理方面的能力。该数据集包含3,026个精心筛选的示例,涵盖21种误导类型和10种图表类型,每个样本均配有标准化图表代码、CSV数据、多选题及标注解释,经多轮MLLM校验与专家人工审核。我们对24个前沿MLLM进行了基准测试,分析其在不同误导类型和图表格式下的表现,并提出一种新型区域感知推理流程,显著提升模型准确率。本工作为开发稳健、可信且符合负责任视觉传播需求的MLLM奠定了基础。

原文摘要 · Abstract (English)

Misleading visualizations, which manipulate chart representations to support specific claims, can distort perception and lead to incorrect conclusions. Despite decades of research, they remain a widespread issue, posing risks to public understanding and raising safety concerns for AI systems involved in data-driven communication. While recent multimodal large language models (MLLMs) show strong chart comprehension abilities, their capacity to detect and interpret misleading charts remains unexplored. We introduce Misleading ChartQA benchmark, a large-scale multimodal dataset designed to evaluate MLLMs on misleading chart reasoning. It contains 3,026 curated examples spanning 21 misleader types and 10 chart types, each with standardized chart code, CSV data, multiple-choice questions, and labeled explanations, validated through iterative MLLM checks and expert human review. We benchmark 24 state-of-the-art MLLMs, analyze their performance across misleader types and chart formats, and propose a novel region-aware reasoning pipeline that enhances model accuracy. Our work lays the foundation for developing MLLMs that are robust, trustworthy, and aligned with the demands of responsible visual communication.

误导图表多模态模型可信AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。