arXiv:2512.21337cs.CV2025-12

发现视觉语言模型对著名建筑有严重偏好,依赖记忆而非理解。

Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models

  • 构建5.5万张建筑图像数据集,带年代、位置和热度标签。
  • 热门建筑预测准确率比普通建筑高34%,暴露记忆依赖问题。
  • 适合关注模型公平性与泛化能力的研究者使用。

我们揭示了当前顶级视觉语言模型存在显著的流行度偏差:在著名建筑上的准确率比普通建筑高出高达34%,表明其更依赖记忆而非可迁移的理解。为系统研究这一现象,我们推出了目前最大的开放基准数据集YearGuessr,包含来自157个国家的55,546张建筑图像,附带连续序数标签(1001–2024年)、地理坐标及页面浏览量(作为流行度代理)。基于该数据集,我们将建筑年代预测任务建模为序数回归,并提出考虑流行度的区间准确率指标以量化偏差。在包含30多个模型(包括我们的YearCLIP)的基准测试中,结果显示模型在热门、被记忆的项目上表现优异,但在不知名对象上显著失准,暴露出其推理能力的关键缺陷。

原文摘要 · Abstract (English)

We expose a significant popularity bias in state-of-the-art vision-language models (VLMs), which achieve up to 34% higher accuracy on famous buildings compared to ordinary ones, indicating a reliance on memorization over generalizable understanding. To systematically investigate this, we introduce the largest open benchmark for this task: the YearGuessr dataset, a collection of 55,546 building images with multi-modal attributes from 157 countries, annotated with continuous ordinal labels of their construction year (1001-2024), GPS data, and page-view counts as a proxy for popularity. Using this dataset, we frame the construction year prediction task as ordinal regression and introduce popularity-aware interval accuracy metrics to quantify this bias. Our resulting benchmark of 30+ models, including our YearCLIP model, confirms that VLMs excel on popular, memorized items but struggle significantly with unrecognized subjects, exposing a critical flaw in their reasoning capabilities. Project page: https://sytwu.github.io/BeyondMemo/

视觉语言模型流行度偏差序数回归多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。