arXiv:2502.20503cs.CL2025-02ACL被引 9

MLLM在误导性图表前易出错,新方法可提升准确率19.6个百分点

Protecting multimodal large language models against misleading visualizations

  • 用表格问答和重绘图表提升模型对误导性图表的识别能力
  • 误导性图表下模型准确率降至随机水平,最高提升19.6个百分点
  • 适用于需可靠图表理解的金融、医疗等高风险场景

可视化在数据驱动时代日益重要。多模态大语言模型(MLLM)在自动图表理解方面进展迅速,但在误导性可视化(扭曲数据、引导错误结论)面前表现脆弱。我们发现,面对误导性图表时,MLLM问答准确率平均降至随机基线水平。为此,我们首次对比六种推理阶段方法,在不损害正常图表表现的前提下提升对误导性图表的鲁棒性。结果表明,基于表格的问答与重绘可视化两种方法有效,准确率最高提升19.6个百分点。代码与数据已公开。

原文摘要 · Abstract (English)

Visualizations play a pivotal role in daily communication in an increasingly data-driven world. Research on multimodal large language models (MLLMs) for automated chart understanding has accelerated massively, with steady improvements on standard benchmarks. However, for MLLMs to be reliable, they must be robust to misleading visualizations, i.e., charts that distort the underlying data, leading readers to draw inaccurate conclusions. Here, we uncover an important vulnerability: MLLM question-answering (QA) accuracy on misleading visualizations drops on average to the level of the random baseline. To address this, we provide the first comparison of six inference-time methods to improve QA performance on misleading visualizations, without compromising accuracy on non-misleading ones. We find that two methods, table-based QA and redrawing the visualization, are effective, with improvements of up to 19.6 percentage points. We make our code and data available.

多模态模型图表理解鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。