arXiv:2508.17398cs.CL2025-08Conference of the …被引 3

首个评估视觉语言模型在交互式数据看板中问答能力的基准

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

  • 构建包含112个真实看板和405个问题的交互式问答数据集
  • 顶尖模型最高准确率仅38.69%,凸显交互推理难度
  • 适合研究人机交互、多模态推理与智能数据分析的学者

看板是支持数据驱动决策的强大可视化工具,整合了多个可交互视图,使用户能够探索、筛选和导航数据。与静态图表不同,看板支持丰富的交互功能,这对发现真实分析工作流中的洞察至关重要。然而,现有的数据可视化问答基准大多忽视这种交互性,仅关注静态图表,严重限制了对现代多模态代理在基于GUI推理方面能力的评估。为填补这一空白,我们提出DashboardQA,这是首个专为评估视觉-语言GUI代理理解与交互真实看板能力而设计的基准。该基准包含来自Tableau Public的112个交互式看板和覆盖五个类别的405个问答对:多项选择、事实型、假设型、跨看板及对话式。通过对多种主流闭源与开源GUI代理进行评估,我们的分析揭示了它们在定位看板元素、规划交互路径和执行推理方面的关键局限。结果表明,所有评估的VLM在交互式看板推理任务上均表现困难。即使表现最佳的Gemini-Pro-2.5代理也仅达到38.69%准确率,而OpenAI CUA代理仅为22.69%,充分证明了该基准的挑战性。我们已在https://github.com/vis-nlp/DashboardQA发布DashboardQA。

原文摘要 · Abstract (English)

Dashboards are powerful visualization tools for data-driven decision-making, integrating multiple interactive views that allow users to explore, filter, and navigate data. Unlike static charts, dashboards support rich interactivity, which is essential for uncovering insights in real-world analytical workflows. However, existing question-answering benchmarks for data visualizations largely overlook this interactivity, focusing instead on static charts. This limitation severely constrains their ability to evaluate the capabilities of modern multimodal agents designed for GUI-based reasoning. To address this gap, we introduce DashboardQA, the first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards. The benchmark includes 112 interactive dashboards from Tableau Public and 405 question-answer pairs with interactive dashboards spanning five categories: multiple-choice, factoid, hypothetical, multi-dashboard, and conversational. By assessing a variety of leading closed- and open-source GUI agents, our analysis reveals their key limitations, particularly in grounding dashboard elements, planning interaction trajectories, and performing reasoning. Our findings indicate that interactive dashboard reasoning is a challenging task overall for all the VLMs evaluated. Even the top-performing agents struggle; for instance, the best agent based on Gemini-Pro-2.5 achieves only 38.69% accuracy, while the OpenAI CUA agent reaches just 22.69%, demonstrating the benchmark's significant difficulty. We release DashboardQA at https://github.com/vis-nlp/DashboardQA

多模态推理交互式看板GUI理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。