arXiv:2508.21675cs.CLcs.CV2025-08ACL被引 5

自动识别误导性图表,防止信息误传。

Is this chart lying to me? Automating the detection of misleading visualizations

  • 构建真实世界图表数据集,标注12类误导设计。
  • 合成5.7万张图表用于模型训练,提升检测能力。
  • 适合关注虚假信息、AI可解释性的研究者使用。

误导性图表是社交媒体和网络上信息错误传播的重要原因。通过违反图表设计原则,它们扭曲数据,导致读者得出错误结论。以往研究表明,人类和多模态大语言模型(MLLMs)常被此类图表欺骗。自动检测误导性图表并识别其违反的设计规则,有助于保护读者并减少信息误传。然而,由于缺乏大规模、多样且公开的数据集,AI模型的训练与评估受到限制。本文提出Misviz基准,包含2,604个真实世界图表,标注了12种误导类型。为支持模型训练,我们还创建了由Matplotlib生成的合成数据集Misviz-synth,共57,665张图表,基于真实数据表。我们在两个数据集上对最先进的MLLMs、基于规则的系统和图像-轴分类器进行了全面评估。结果表明,该任务仍极具挑战性。我们开源了Misviz、Misviz-synth及配套代码。

原文摘要 · Abstract (English)

Misleading visualizations are a potent driver of misinformation on social media and the web. By violating chart design principles, they distort data and lead readers to draw inaccurate conclusions. Prior work has shown that both humans and multimodal large language models (MLLMs) are frequently deceived by such visualizations. Automatically detecting misleading visualizations and identifying the specific design rules they violate could help protect readers and reduce the spread of misinformation. However, the training and evaluation of AI models has been limited by the absence of large, diverse, and openly available datasets. In this work, we introduce Misviz, a benchmark of 2,604 real-world visualizations annotated with 12 types of misleaders. To support model training, we also create Misviz-synth, a synthetic dataset of 57,665 visualizations generated using Matplotlib and based on real-world data tables. We perform a comprehensive evaluation on both datasets using state-of-the-art MLLMs, rule-based systems, and image-axis classifiers. Our results reveal that the task remains highly challenging. We release Misviz, Misviz-synth, and the accompanying code.

图表检测误导信息AI安全数据可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。