arXiv:2608.10567cs.AI2026-08

首个评估大模型生成交互式数据看板的基准,强调可重现的分析流程。

DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation

论文配图:DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation
图 1 · 摘自论文原文
  • 要求模型输出看板与可重放的操作轨迹,实现动态评估。
  • 引入视觉语言模型裁判,基于操作证据打分,提升评价一致性。
  • 适合关注交互式生成与真实可用性的研究人员。

分析型看板通过协调的视图和交互支持数据探索与决策。近期模型能从数据和自然语言目标生成看板,但评估其实际效用仍具挑战。看板生成具有开放性,仅考察静态外观或执行成功不足以反映分析支持与交互质量。我们提出 DashArena,据我们所知首个面向开放性、任务驱动的交互式分析看板生成基准。其核心创新在于要求系统同时生成看板和可重放的交互轨迹。浏览器执行器回放轨迹,将系统的预期分析流程转化为可复现的视觉与执行证据。视觉语言模型裁判利用该证据比较候选方案,结合 Bradley-Terry 模型生成排行榜。我们进一步将裁判蒸馏为开源的 DashJudge-8B。人类评估表明,DashJudge-8B 能有效复现人工判断;消融实验显示,交互证据显著提升裁判一致性。对前沿模型的实验揭示了持续存在的渲染、分析与交互失败问题。这些结果表明,真实看板生成仍具挑战,且交互感知评估能捕捉静态或仅执行检查遗漏的缺陷。

原文摘要 · Abstract (English)

Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and natural-language goals, but evaluating their usefulness remains difficult. Dashboard generation is open-ended, and neither static appearance nor successful execution alone captures analytical support and interaction quality. We introduce DashArena, to our knowledge the first benchmark for open-ended, task-grounded generation of interactive analytic dashboards. Its key innovation is to require each system to generate both a dashboard and a replayable interaction trajectory. A browser executor replays the trajectory and turns the system's intended analytical workflow into reproducible visual and execution evidence. A VLM judge compares candidates using this evidence, and Bradley--Terry aggregation produces the leaderboard. We further distill the judge into the open-weight DashJudge-8B. Human evaluations show that DashJudge-8B effectively reproduces human judgments and ablations show that interaction evidence improves judge agreement. Experiments with frontier models reveal persistent rendering, analytical, and interaction failures. Together, these results show that realistic dashboard generation remains challenging and that interaction-aware evaluation captures failures missed by static or execution-only checks.

大模型评估交互生成看板生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。