arXiv:2510.22373cs.CLcs.AI2025-10中稿 · ICLR被引 17

首个可视化美学与质量评估基准,揭示大模型短板并提出针对性改进方案。

VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations

  • 构建3090个真实场景的可视化样本库,涵盖32种图表类型。
  • 顶尖大模型在评估上与人类专家仍有显著差距,误差达0.553。
  • 提出VisJudge模型,性能超越GPT-5,误差降低23.9%。

可视化是将复杂数据转化为直观洞察的有效方式,其价值取决于数据表达准确性、信息传达清晰度和视觉美感。然而,评估可视化质量极具挑战:需同时判断数据编码、信息表现力与审美设计。尽管多模态大语言模型(MLLM)在自然图像审美评估中表现优异,但尚无系统性基准衡量其在可视化评估中的能力。为此,我们提出VisJudge-Bench,首个全面评估MLLM在可视化美学与质量评估表现的基准。该基准包含3,090个来自真实场景的专家标注样本,覆盖单图、多图及仪表板,涵盖32种图表类型。系统测试显示,即使最先进的模型(如GPT-5)与人类专家仍存在显著差距,平均绝对误差(MAE)为0.553,与人类评分相关性仅为0.428。为解决此问题,我们提出VisJudge模型,专门用于可视化美学与质量评估。实验表明,该模型显著缩小与人类判断的差距,将MAE降至0.421(下降23.9%),与人类一致性提升至0.687(提高60.5%)。基准代码已开源:https://github.com/HKUSTDial/VisJudgeBench。

原文摘要 · Abstract (English)

Visualization, a domain-specific yet widely used form of imagery, is an effective way to turn complex datasets into intuitive insights, and its value depends on whether data are faithfully represented, clearly communicated, and aesthetically designed. However, evaluating visualization quality is challenging: unlike natural images, it requires simultaneous judgment across data encoding accuracy, information expressiveness, and visual aesthetics. Although multimodal large language models (MLLMs) have shown promising performance in aesthetic assessment of natural images, no systematic benchmark exists for measuring their capabilities in evaluating visualizations. To address this, we propose VisJudge-Bench, the first comprehensive benchmark for evaluating MLLMs' performance in assessing visualization aesthetics and quality. It contains 3,090 expert-annotated samples from real-world scenarios, covering single visualizations, multiple visualizations, and dashboards across 32 chart types. Systematic testing on this benchmark reveals that even the most advanced MLLMs (such as GPT-5) still exhibit significant gaps compared to human experts in judgment, with a Mean Absolute Error (MAE) of 0.553 and a correlation with human ratings of only 0.428. To address this issue, we propose VisJudge, a model specifically designed for visualization aesthetics and quality assessment. Experimental results demonstrate that VisJudge significantly narrows the gap with human judgment, reducing the MAE to 0.421 (a 23.9% reduction) and increasing the consistency with human experts to 0.687 (a 60.5% improvement) compared to GPT-5. The benchmark is available at https://github.com/HKUSTDial/VisJudgeBench.

可视化评估大模型评测MLLM基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。