构建多模态知识更新基准,评估模型对变化知识的适应能力。
MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge
- 设计涵盖2.5万+知识实例的多模态更新测试集
- 发现微调易导致灾难性遗忘,知识编辑更擅长持续更新
- 适合关注多模态模型知识演进的研究者使用
随着现实世界知识不断演变,多模态模型在预训练中获取的参数化知识越来越难以与真实世界知识保持一致。现有研究仅关注学习此前未知的知识,忽略了对模型已掌握但后续发生变化的知识进行更新;此外,评估局限于单一模态,缺乏跨模态一致性分析。为此,本文提出MMKU-Bench,一个全面的多模态知识更新评估基准,包含超过25,000个知识实例和49,000多张图像,覆盖已更新知识与未知知识两种场景,支持不同类型知识学习的对比分析。在此基准上,我们评估了多种代表性方法,包括监督微调(SFT)、基于人类反馈的强化学习(RLHF)和知识编辑(KE)。实验结果表明,SFT和RLHF容易引发灾难性遗忘,而知识编辑虽更好保留通用能力,但在持续更新方面存在明显局限。总体而言,MMKU-Bench为多模态知识更新提供了可靠且全面的评估基础,推动该领域发展。
原文摘要 · Abstract (English)
As real-world knowledge continues to evolve, the parametric knowledge acquired by multimodal models during pretraining becomes increasingly difficult to remain consistent with real-world knowledge. Existing research on multimodal knowledge updating focuses only on learning previously unknown knowledge, while overlooking the need to update knowledge that the model has already mastered but that later changes; moreover, evaluation is limited to the same modality, lacking a systematic analysis of cross-modal consistency. To address these issues, this paper proposes MMKU-Bench, a comprehensive evaluation benchmark for multimodal knowledge updating, which contains over 25k knowledge instances and more than 49k images, covering two scenarios, updated knowledge and unknown knowledge, thereby enabling comparative analysis of learning across different knowledge types. On this benchmark, we evaluate a variety of representative approaches, including supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and knowledge editing (KE). Experimental results show that SFT and RLHF are prone to catastrophic forgetting, while KE better preserve general capabilities but exhibit clear limitations in continual updating. Overall, MMKU-Bench provides a reliable and comprehensive evaluation benchmark for multimodal knowledge updating, advancing progress in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。