arXiv:2604.07338cs.CVcs.CL2026-04

构建跨文化图像元数据推理基准,揭示视觉语言模型在文化理解上的局限。

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

  • 提出多类别跨文化基准,评估模型从图像推断创作者、起源等结构化信息的能力。
  • 模型在不同文化区域表现差异大,属性准确率最高仅达63.2%,且多数结果不完整。
  • 适合关注文化智能、数字人文与多模态模型公平性的研究者参考。

近年来,视觉语言模型(VLMs)在文化遗产图像描述方面取得进展,但基于视觉输入推断结构化文化元数据(如创作者、起源、时期)仍缺乏深入探索。本文提出一个多类别、跨文化的基准任务,采用大语言模型作为评判者(LLM-as-Judge)框架,衡量模型输出与参考标注的语义一致性。为评估文化推理能力,报告了跨文化区域的精确匹配、部分匹配及属性级准确率。结果显示,模型仅捕捉到碎片化信号,在不同文化和元数据类型间表现差异显著,预测结果不一致且支撑薄弱。这些发现凸显了当前VLMs在超越视觉感知的文化元数据推断中的局限性。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, period) from visual input remains underexplored. We introduce a multi-category, cross-cultural benchmark for this task and evaluate VLMs using an LLM-as-Judge framework that measures semantic alignment with reference annotations. To assess cultural reasoning, we report exact-match, partial-match, and attribute-level accuracy across cultural regions. Results show that models capture fragmented signals and exhibit substantial performance variation across cultures and metadata types, leading to inconsistent and weakly grounded predictions. These findings highlight the limitations of current VLMs in structured cultural metadata inference beyond visual perception.

文化推理多模态元数据跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。