arXiv:2411.12790cs.CVcs.AI2024-11ICCV被引 14

让多模态大模型精准修改图像中多个实体的知识

Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models

  • 用视觉与文本联合判断编辑范围,实现细粒度精准更新
  • 在新构建的FGVEdit基准上超越现有方法
  • 适合需要精确修正图像知识的研究者和开发者

知识编辑旨在高效低成本地修正错误信息和更新过时内容。近期研究将知识编辑从纯语言模型扩展至融合文本与视觉信息的多模态大模型(MLLM),引入了新的编辑挑战。现有工作主要聚焦于文本主导、粗粒度场景,难以应对多模态上下文的独特需求。本文提出一种面向视觉的细粒度多模态知识编辑任务,针对图像中多个交互实体进行精确编辑。为此,我们构建了细粒度视觉知识编辑(FGVEdit)基准。同时提出基于多模态作用域分类器的知识编辑框架(MSCKE),该框架融合视觉与文本信息,准确识别并更新图像中特定实体的相关知识,确保无关信息不受影响,克服传统仅依赖文本编辑的局限性。在FGVEdit基准上的大量实验表明,MSCKE优于现有方法,有效解决了多模态知识编辑的复杂挑战。

原文摘要 · Abstract (English)

Knowledge editing aims to efficiently and cost-effectively correct inaccuracies and update outdated information. Recently, there has been growing interest in extending knowledge editing from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs), which integrate both textual and visual information, introducing additional editing complexities. Existing multimodal knowledge editing works primarily focus on text-oriented, coarse-grained scenarios, failing to address the unique challenges posed by multimodal contexts. In this paper, we propose a visual-oriented, fine-grained multimodal knowledge editing task that targets precise editing in images with multiple interacting entities. We introduce the Fine-Grained Visual Knowledge Editing (FGVEdit) benchmark to evaluate this task. Moreover, we propose a Multimodal Scope Classifier-based Knowledge Editor (MSCKE) framework. MSCKE leverages a multimodal scope classifier that integrates both visual and textual information to accurately identify and update knowledge related to specific entities within images. This approach ensures precise editing while preserving irrelevant information, overcoming the limitations of traditional text-only editing methods. Extensive experiments on the FGVEdit benchmark demonstrate that MSCKE outperforms existing methods, showcasing its effectiveness in solving the complex challenges of multimodal knowledge editing.

多模态知识编辑视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。