arXiv:2606.08959cs.CVcs.CL2026-06

构建中国世遗文化理解数据集,评估模型对文化遗产的深层认知能力

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

论文配图:ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China
图 1 · 摘自论文原文
  • 基于联合国教科文组织标准构建多模态问答数据集
  • 顶尖模型在视觉识别上超人类,但文化推理能力仍不足
  • 适合研究跨文化多模态学习与文化遗产智能理解的学者

我们提出ChinaHeritaQA,一个用于评估视觉语言模型(VLMs)在中国世界遗产地文化推理能力的多模态基准数据集。该数据集包含2,279张真实场景图像,对应14,133个中英文双语多项选择题,涵盖从身份识别到历史分期、建筑分析等七类认知维度。数据集基于联合国教科文组织遗产本体构建,并经严格人工标注,确保语言质量与事实一致性。对前沿VLMs的评估显示,尽管顶级模型平均表现超过人类,但在任务层面存在显著差异:模型在视觉识别上表现优异,却在文化语境推理上明显受限,且性能随朝代与地区不同而变化。结果表明,强大的视觉检索能力无法直接转化为文化与历史理解能力。数据集已公开,以支持未来文化感知多模态学习研究。

原文摘要 · Abstract (English)

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis. Guided by a UNESCO-aligned heritage ontology and verified through rigorous human annotation, the dataset ensures linguistic quality and factual consistency. Evaluations of state-of-the-art VLMs reveal that while top models exceed human performance on average, substantial task-level variation emerges: models excel at visual recognition but struggle with culturally grounded reasoning. Performance also varies by dynasty and region. ChinaHeritaQA reveals that strong visual retrieval does not extend to cultural and historical understanding. We release the dataset to support future research on culturally aware multimodal learning.

文化理解多模态世遗保护视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。