arXiv:2510.19457cs.CL2025-10ACL被引 6

构建首个多模态时序知识评估基准,检验大模型对动态事实的理解能力。

MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models

  • 设计6维度11任务的时序知识评测框架,覆盖认知、推理与鲁棒性等
  • 15个主流大模型在该基准上平均最佳得分仅63.07(Gemini-2.5-Pro)
  • 发现开源模型普遍缺乏时序理解能力,且体育类知识最弱

大型多模态模型(LMMs)通过跨模态预训练编码丰富事实知识,但其静态表征难以保持对时序敏感知识的准确理解。现有基准受限于静态设计,无法充分评估LMMs对时序知识的掌握能力。为此,我们提出MINED,一个全面的评估基准,涵盖6个关键维度和11项挑战性任务:认知、意识、可信度、理解、推理与鲁棒性。MINED基于维基百科由两名专业标注员构建,包含2,104个时序敏感知识样本,覆盖六类知识类型。在该基准上对15个广泛使用的LMMs进行评估显示,Gemini-2.5-Pro达到最高平均CEM分数63.07,而大多数开源LMM仍缺乏时序理解能力。模型在组织类知识表现最好,体育类知识最差。为进一步应对挑战,我们探究了通过知识编辑方法更新时序知识的可行性,结果表明在单次编辑场景下,LMM能有效完成知识更新。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) encode rich factual knowledge via cross-modal pre-training, yet their static representations struggle to maintain an accurate understanding of time-sensitive factual knowledge. Existing benchmarks remain constrained by static designs, inadequately evaluating LMMs' ability to understand time-sensitive knowledge. To address this gap, we propose MINED, a comprehensive benchmark that evaluates temporal awareness along 6 key dimensions and 11 challenging tasks: cognition, awareness, trustworthiness, understanding, reasoning, and robustness. MINED is constructed from Wikipedia by two professional annotators, containing 2,104 time-sensitive knowledge samples spanning six knowledge types. Evaluating 15 widely used LMMs on MINED shows that Gemini-2.5-Pro achieves the highest average CEM score of 63.07, while most open-source LMMs still lack time understanding ability. Meanwhile, LMMs perform best on organization knowledge, whereas their performance is weakest on sport. To address these challenges, we investigate the feasibility of updating time-sensitive knowledge in LMMs through knowledge editing methods and observe that LMMs can effectively update knowledge via knowledge editing methods in single editing scenarios.

多模态时序知识知识编辑评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。