评测大模型对图表体验影响的预测能力,发现其在对比判断上更可靠。
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
- 构建36张图表的体验影响评估数据集,涵盖7种体验维度。
- 模型对单张图表判断不够敏感,但在图表对比中准确率高。
- 适合关注人机体验评估差异的研究者和产品设计人员。
多模态大模型(MLLM)在视觉理解任务中取得显著进展,为预测图表的感知与情感影响提供了巨大潜力。然而,许多应用基于少数样本的过度泛化假设,缺乏充分验证。本文提出Chart-to-Experience基准数据集,包含36张图表,由众包工作者评估其在七种体验因素上的影响。以该数据集为真实标签,我们评估了先进MLLM在两项任务中的表现:直接预测与图表成对比较。结果表明,MLLM在评估单张图表时不如人类敏感,但在成对比较中表现出较高准确性和可靠性。
原文摘要 · Abstract (English)
The field of Multimodal Large Language Models (MLLMs) has made remarkable progress in visual understanding tasks, presenting a vast opportunity to predict the perceptual and emotional impact of charts. However, it also raises concerns, as many applications of LLMs are based on overgeneralized assumptions from a few examples, lacking sufficient validation of their performance and effectiveness. We introduce Chart-to-Experience, a benchmark dataset comprising 36 charts, evaluated by crowdsourced workers for their impact on seven experiential factors. Using the dataset as ground truth, we evaluated capabilities of state-of-the-art MLLMs on two tasks: direct prediction and pairwise comparison of charts. Our findings imply that MLLMs are not as sensitive as human evaluators when assessing individual charts, but are accurate and reliable in pairwise comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。