测试视觉语言模型对随时间变化事实的掌握能力,发现模型常输出过时信息。
V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models
- 构建动态基准V-DyKnow,评估模型在时间敏感事实上的表现
- 模型对图像输入的事实可靠性低于文本,且更新知识效果差
- 适合关注模型时效性、知识更新与多模态对齐的研究者
视觉语言模型(VLM)在静态文档快照(含图像与文本)上训练,其训练数据和评测基准通常固定,隐含将事实知识视为不变。但真实世界事实具有时间敏感性,会随时间发生突发或周期性变化,导致模型预测过时。本文提出V-DyKnow,一个用于评估VLM中时间敏感事实知识的视觉动态知识基准。通过该基准,我们评测了闭源与开源VLM,并分析:a) 模型在跨模态及输入扰动下的响应可靠性(正确性与一致性);b) 知识编辑与多模态RAG方法在跨模态知识更新中的有效性;c) 过时预测的来源,结合数据与机制分析。结果表明,模型频繁输出过时事实,反映出训练阶段使用了过时的数据快照。事实可靠性从文本向图像递减,即使实体识别正确亦然。此外,现有对齐方法无法一致地跨模态更新知识。这些发现揭示了当前VLM在跨模态获取与更新时间敏感知识方面的根本局限。我们已发布基准、代码与评估数据。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are trained on data snapshots of documents, including images and texts. Their training data and evaluation benchmarks are typically static, implicitly treating factual knowledge as time-invariant. However, real-world facts are intrinsically time-sensitive and subject to erratic and periodic changes, causing model predictions to become outdated. We present V-DyKnow, a Visual Dynamic Knowledge benchmark for evaluating time-sensitive factual knowledge in VLMs. Using V-DyKnow, we benchmark closed- and open-source VLMs and analyze a) the reliability (correctness and consistency) of model responses across modalities and input perturbations; b) the efficacy of knowledge editing and multi-modal RAG methods for knowledge updates across modalities; and c) the sources of outdated predictions, through data and mechanistic analysis. Our results show that VLMs frequently output outdated facts, reflecting outdated snapshots used in the (pre-)training phase. Factual reliability degrades from textual to visual stimuli, even when entities are correctly recognized. Besides, existing alignment approaches fail to consistently update the models' knowledge across modalities. Together, these findings highlight fundamental limitations in how current VLMs acquire and update time-sensitive knowledge across modalities. We release the benchmark, code, and evaluation data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。