arXiv:2505.24449cs.CL2025-05被引 6

测试大模型对动态多模态知识的更新能力,发现旧知识易被覆盖。

When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations

  • 构建新基准MMEVOKE,评估多模态动态知识注入能力
  • 现有方法知识注入效果差,且会引发模型能力退化
  • 引入知识增强与数据回放策略,有效提升注入效果并保护旧知识

大型多模态模型(LMMs)虽存储大量预训练知识,但难以跟上现实世界的知识更新,导致获取动态知识时能力下降。当前多数研究聚焦静态文本知识注入,忽视了动态多模态知识注入,使LMM在多模态知识注入方面的潜力成谜。为此,我们提出构建MMEVOKE基准,用于评估LMM在多模态演化知识注入中的表现,该基准包含9,422个样本,覆盖159种子类型。基于在MMEVOKE上的大量实验,我们揭示现有知识注入方法存在注入效果差和通用能力退化等问题。为应对挑战,我们引入知识增强与知识保留方法,发现知识感知增强可提升注入性能,而数据回放(Data Replay)与混合专家(MoE)方法能有效缓解能力退化。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) store vast amounts of pretrained knowledge but struggle to remain aligned with real-world updates, making it difficult to avoid capability degradation when acquiring evolving knowledge. Furthermore, most current work focuses on exploring static textual knowledge injection, neglecting dynamic multimodal evolving knowledge injection, leaving the potential of LMMs for multimodal knowledge injection as an open question. To address this, we first propose a pipeline to construct MMEVOKE, a benchmark for evaluating LMMs' ability in multimodal evolving knowledge injection. MMEVOKE contains 9,422 samples spanning 159 subtypes. Then, based on extensive experiments with MMEVOKE, we reveal challenges such as poor injection performance and capability degradation in existing knowledge injection methods through knowledge injection tests and general capability tests. Finally, to tackle these challenges, we introduce knowledge augmentation and knowledge retention methods, finding that knowledge-aware augmentation strengthens knowledge injection performance, and that Data Replay and MoE methods effectively mitigate capability degradation.

多模态模型知识注入持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。