arXiv:2505.17235cs.CVcs.CL2025-05被引 1

评测大模型在噪声图表下的理解能力,发现现有模型易受干扰。

CHAOS: Chart Analysis with Outlier Samples

  • 构建含15类扰动的基准测试,覆盖文本与视觉问题。
  • 3个难度等级下,主流模型在图表问答任务中准确率下降超40%。
  • 适合关注图表理解鲁棒性的研究人员和应用开发者。

图表在数据分析与可视化中至关重要,但现实应用中常出现噪声或异常特征。本文提出CHAOS(CHart Analysis with Outlier Samples),一个系统性评估多模态大模型在图表扰动下表现的鲁棒性基准。CHAOS包含5种文本扰动和10种视觉扰动,每类扰动设置三个严重程度(易、中、难),参考人类评估结果设计。基准涵盖13个先进多模态大模型,按训练范围分为通用型、文档型和图表专用型三类。实验基于两个下游任务(ChartQA和Chart-to-Text)进行,全面分析模型在不同扰动下的表现。案例研究揭示了模型对各类扰动的脆弱性,为未来图表理解研究提供方向。数据与代码已公开于:http://huggingface.co/datasets/omoured/CHAOS。

原文摘要 · Abstract (English)

Charts play a critical role in data analysis and visualization, yet real-world applications often present charts with challenging or noisy features. However, "outlier charts" pose a substantial challenge even for Multimodal Large Language Models (MLLMs), which can struggle to interpret perturbed charts. In this work, we introduce CHAOS (CHart Analysis with Outlier Samples), a robustness benchmark to systematically evaluate MLLMs against chart perturbations. CHAOS encompasses five types of textual and ten types of visual perturbations, each presented at three levels of severity (easy, mid, hard) inspired by the study result of human evaluation. The benchmark includes 13 state-of-the-art MLLMs divided into three groups (i.e., general-, document-, and chart-specific models) according to the training scope and data. Comprehensive analysis involves two downstream tasks (ChartQA and Chart-to-Text). Extensive experiments and case studies highlight critical insights into robustness of models across chart perturbations, aiming to guide future research in chart understanding domain. Data and code are publicly available at: http://huggingface.co/datasets/omoured/CHAOS.

图表理解大模型评测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。