arXiv:2503.04095cs.CLcs.AI2025-03被引 3

让大模型学会看图做假设推理,突破死记硬背式回答

Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts

  • 设计假设性问题,逼模型基于图表内容做反事实推理
  • 18个大模型在新任务上表现普遍不佳,推理能力不均衡
  • 用人工+AI协作生成数据,低成本构建高质量评测集

多模态大语言模型(MLLMs)在视觉-语义理解方面备受关注。现有图表评测基准主要考察模型解析图表信息回答问题的能力,但忽略了模型依赖参数记忆而非真正理解图表内容的固有偏差。为解决这一问题,本文提出一种新的图表假设性问答(HQA)任务,通过设定相同问题的不同假设条件,迫使模型基于图表内容进行反事实推理。同时,提出一种人机协同的数据合成方法HAI,利用大模型高效的文本编辑能力与人类专家知识,以低成本生成多样且高质量的HQA数据。基于公开数据源,我们构建了名为Chart-HQA的挑战性评测基准。对18种不同规模的MLLMs评估结果显示,当前模型在该任务上存在显著泛化困难,且推理表现严重不平衡。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have garnered significant attention for their strong visual-semantic understanding. Most existing chart benchmarks evaluate MLLMs' ability to parse information from charts to answer questions. However, they overlook the inherent output biases of MLLMs, where models rely on their parametric memory to answer questions rather than genuinely understanding the chart content. To address this limitation, we introduce a novel Chart Hypothetical Question Answering (HQA) task, which imposes assumptions on the same question to compel models to engage in counterfactual reasoning based on the chart content. Furthermore, we introduce HAI, a human-AI interactive data synthesis approach that leverages the efficient text-editing capabilities of LLMs alongside human expert knowledge to generate diverse and high-quality HQA data at a low cost. Using HAI, we construct Chart-HQA, a challenging benchmark synthesized from publicly available data sources. Evaluation results on 18 MLLMs of varying model sizes reveal that current models face significant generalization challenges and exhibit imbalanced reasoning performance on the HQA task.

图表理解假设推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。