arXiv:2601.07984cs.CL2026-01被引 1

首个跨文化艺术评述评估框架,揭示VLM在非西方艺术上评分偏高

Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models

  • 构建三级评估体系,融合艺术理论与专家评分校准
  • 15个VLM在294对作品中表现:非西方艺术得分普遍偏低
  • 发现自动指标与人工评分衡量不同维度,单评委更可靠

视觉语言模型(VLMs)在图像描述上表现优异,但在文化解读方面缺乏有效评估。现有基准仅评测感知能力,常用代理指标如自动评分和大模型评分平均值在文化敏感性生成任务中不可靠。本文提出基于艺术理论的三级评估框架(第2节),通过五个层级(L1–L5)和165个文化特定维度,覆盖六种艺术传统:一级计算自动化质量指标,二级采用评分表进行单评委打分,三级通过逻辑斯蒂校准将聚合分数与人类专家评分对齐。在15个VLM、294组评估对上的应用表明:(i) 自动化指标与评委打分衡量不同构念,单评委校准更可靠;(ii) 文化理解能力从视觉描述(L1–L2)到文化阐释(L3–L5)逐级下降;(iii) 西方艺术样本得分始终高于非西方样本。据我们所知,这是首个用于生成式艺术评述的跨文化评估工具,提供可复现的VLM文化能力审计方法。框架代码已公开于https://github.com/yha9806/VULCA-Framework。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel at visual description yet remain under-validated for cultural interpretation. Existing benchmarks assess perception without interpretation, and common evaluation proxies, such as automated metrics and LLM-judge averaging, are unreliable for culturally sensitive generative tasks. We address this measurement gap with a tri-tier evaluation framework grounded in art-theoretical constructs (Section 2). The framework operationalises cultural understanding through five levels (L1--L5) and 165 culture-specific dimensions across six traditions: Tier I computes automated quality indicators, Tier II applies rubric-based single-judge scoring, and Tier III calibrates the aggregate score to human expert ratings via sigmoid calibration. Applied to 15 VLMs across 294 evaluation pairs, the validated instrument reveals that (i) automated metrics and judge scoring measure different constructs, establishing single-judge calibration as the more reliable alternative; (ii) cultural understanding degrades from visual description (L1--L2) to cultural interpretation (L3--L5); and (iii) Western art samples consistently receive higher scores than non-Western ones. To our knowledge, this is the first cross-cultural evaluation instrument for generative art critique, providing a reproducible methodology for auditing VLM cultural competence. Framework code is available at https://github.com/yha9806/VULCA-Framework.

艺术生成跨文化评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。