arXiv:2502.19870cs.CL2025-02ICLR被引 26

构建多模态知识编辑新基准,评估模型对真实场景视觉知识的修改能力。

MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge

  • 设计三类编辑任务:实体、语义与用户定制化,覆盖真实场景复杂性
  • 包含2940条知识与8363张图像,跨33个类别自动生成并人工验证评测题
  • 揭示当前方法在视觉和个性化编辑上仍存挑战,推动领域发展

知识编辑技术已成为更新大语言模型(LLMs)和多模态模型(LMMs)事实知识的关键工具,可在无需从头训练的情况下修正过时或错误信息。然而,现有多模态知识编辑基准主要聚焦于以简单三元组表示的实体级知识,难以捕捉现实世界多模态信息的复杂性。为此,我们提出MMKE-Bench,一个全面的多模态知识编辑基准,旨在评估LMMs在真实场景中编辑多样化视觉知识的能力。该基准引入三类编辑任务:视觉实体编辑、视觉语义编辑和用户特定编辑,并采用自由形式自然语言表示与编辑知识,提供更灵活高效的格式。基准包含2,940条知识与8,363张图像,覆盖33个广泛类别,评测问题自动生成并经人工验证。我们在三个主流LMMs上评估了五种前沿知识编辑方法,结果表明无一方法在所有指标上均占优,且视觉与用户特定编辑尤为困难。MMKE-Bench为评估多模态知识编辑技术的鲁棒性设立了新标准,推动该快速演进领域的进步。

原文摘要 · Abstract (English)

Knowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information without retraining from scratch. However, existing benchmarks for multimodal knowledge editing primarily focus on entity-level knowledge represented as simple triplets, which fail to capture the complexity of real-world multimodal information. To address this issue, we introduce MMKE-Bench, a comprehensive MultiModal Knowledge Editing Benchmark, designed to evaluate the ability of LMMs to edit diverse visual knowledge in real-world scenarios. MMKE-Bench addresses these limitations by incorporating three types of editing tasks: visual entity editing, visual semantic editing, and user-specific editing. Besides, MMKE-Bench uses free-form natural language to represent and edit knowledge, offering a more flexible and effective format. The benchmark consists of 2,940 pieces of knowledge and 8,363 images across 33 broad categories, with evaluation questions automatically generated and human-verified. We assess five state-of-the-art knowledge editing methods on three prominent LMMs, revealing that no method excels across all criteria, and that visual and user-specific edits are particularly challenging. MMKE-Bench sets a new standard for evaluating the robustness of multimodal knowledge editing techniques, driving progress in this rapidly evolving field.

多模态知识编辑评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。