arXiv:2605.15071cs.CVcs.AI2026-05ACL

发现视觉语言模型在解读历史文物时存在时代错位问题,影响文化遗产理解。

On the Cultural Anachronism and Temporal Reasoning in Vision Language Models

论文配图:On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
图 1 · 摘自论文原文
  • 构建了包含600个问题的时序错位评测集(TAB-VLM),覆盖1600件印度文物。
  • 十款主流模型平均准确率仅58.7%,最佳模型仍不足六成。
  • 揭示非西方文化在训练数据中的缺失导致模型难以正确理解历史语境。

视觉语言模型(VLMs)被广泛应用于文化遗产材料,如数字档案与教育平台。本文识别出模型在解读历史文物时存在的根本性问题:文化时代错位,即使用时间上不恰当的概念、材料或文化框架误解历史物件。为量化该现象,我们提出视觉语言模型时序错位基准测试(TAB-VLM),包含600个跨六个类别的问题,评估1600件从史前到现代的印度文化文物的时序推理能力。对十款先进模型的系统评估显示显著缺陷,即使最优模型GPT-5.2也仅达58.7%整体准确率。性能差距在不同架构与规模下持续存在,表明文化时代错位是视觉AI系统普遍存在的局限,与模型大小无关。研究凸显当前VLM能力与文化遗产精准解读需求之间的鸿沟,尤其针对训练数据中代表性不足的非西方视觉文化。本基准为提升多模态AI在历史文物交互中的时序认知能力奠定基础。数据集与代码已公开。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work identifies a fundamental issue in how these models interpret historical artifacts. We define this phenomenon as cultural anachronism, the tendency to misinterpret historical objects using temporally inappropriate concepts, materials, or cultural frameworks. To quantify this phenomenon, we introduce the Temporal Anachronism Benchmark for Vision-Language Models (TAB-VLM), a dataset of 600 questions across six categories, designed to evaluate temporal reasoning on 1,600 Indian cultural artifacts spanning prehistoric to modern periods. Systematic evaluations of ten state-of-the-art models reveal significant deficiencies on our benchmark, and even the best model (GPT-5.2) achieves only 58.7% overall accuracy. The performance gap persists across varying architectures and scales, suggesting that cultural anachronism represents a significant limitation in visual AI systems, regardless of model size. These findings highlight the disparity between current VLM capabilities and the requirements for accurately interpreting cultural heritage materials, particularly for non-Western visual cultures underrepresented in training data. Our benchmark provides a foundation for enhancing temporal cognition in multimodal AI systems that interact with historical artifacts. The dataset and code are available in our project page.

视觉语言模型文化理解时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。