arXiv:2504.04221cs.CV2025-04被引 1

用经典图表实验测试大模型看图能力,发现它有时比人准,有时却差。

Evaluating Graphical Perception with Multimodal LLMs

  • 复现1984年经典图表感知实验,对比大模型与人类表现
  • 部分任务中大模型误差低于人类,但整体表现不一致
  • 揭示视觉分析中模型的优势与局限,适合研究人机协作

多模态大语言模型(MLLM)在图像理解方面取得了显著进展,但在图表中精确回归数值方面仍研究不足。本文通过复现Cleveland和McGill于1984年的经典实验,评估MLLM在图形感知任务中的表现,并与人类表现进行对比。研究重点考察微调和预训练模型在零样本提示下的表现,以判断其是否接近人类对图形信息的感知能力。结果表明,在某些任务中MLLM的表现优于人类,但在其他任务中则表现较差。研究系统呈现各项实验结果,旨在厘清大模型在数据可视化应用中的优势与短板。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have remarkably progressed in analyzing and understanding images. Despite these advancements, accurately regressing values in charts remains an underexplored area for MLLMs. For visualization, how do MLLMs perform when applied to graphical perception tasks? Our paper investigates this question by reproducing Cleveland and McGill's seminal 1984 experiment and comparing it against human task performance. Our study primarily evaluates fine-tuned and pretrained models and zero-shot prompting to determine if they closely match human graphical perception. Our findings highlight that MLLMs outperform human task performance in some cases but not in others. We highlight the results of all experiments to foster an understanding of where MLLMs succeed and fail when applied to data visualization.

图表理解多模态认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。