构建多文化视觉语言评估基准,测试模型对深层文化意义的理解能力。
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
- 设计五层文化理解框架,从视觉感知到哲学美学逐层递进。
- 包含7410对跨八种文化的图文评论对,覆盖中英双语。
- 发现高阶文化推理比基础视觉分析难得多,适合评估模型深层理解力。
我们提出VULCA-Bench,一个用于评估视觉语言模型(VLMs)在超越表层视觉感知基础上的文化理解能力的多文化艺术评论基准。现有VLM基准主要衡量对象识别、场景描述和事实问答等低层次能力(L1-L2),而忽视更高阶的文化解读。VULCA-Bench包含7,410组匹配的图像-评论对,涵盖八种文化传统,并支持中英双语。我们基于五层框架(L1-L5,从视觉感知到哲学美学)定义文化理解,共衍生225个文化特定维度,并由专家撰写双语评论作为支撑。初步实验表明,高层级推理(L3-L5)始终比视觉与技术分析(L1-L2)更具挑战性。数据集、评估脚本和标注工具已开源,许可协议为CC BY 4.0,详见https://github.com/yha9806/VULCA-Bench。
原文摘要 · Abstract (English)
We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models' (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure L1-L2 capabilities (object recognition, scene description, and factual question answering) while under-evaluate higher-order cultural interpretation. VULCA-Bench contains 7,410 matched image-critique pairs spanning eight cultural traditions, with Chinese-English bilingual coverage. We operationalise cultural understanding using a five-layer framework (L1-L5, from Visual Perception to Philosophical Aesthetics), instantiated as 225 culture-specific dimensions and supported by expert-written bilingual critiques. Our pilot results indicate that higher-layer reasoning (L3-L5) is consistently more challenging than visual and technical analysis (L1-L2). The dataset, evaluation scripts, and annotation tools are available under CC BY 4.0 at https://github.com/yha9806/VULCA-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。