arXiv:2412.12821cs.CV2024-12中稿 · AAAI被引 7

构建多模态知识编辑综合评测框架,解决旧信息难修正难题

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing

  • 设计八类任务覆盖多数据集,全面评估编辑效果
  • 提出KGI/KPI双指标,不依赖人工合成样本来测泛化与保留能力
  • 给出分层上下文编辑基线方法,兼顾各项指标表现

大型多模态语言模型(MLLMs)虽在自然语言处理与视觉理解上取得突破,但常包含过时或错误信息。现有多模态知识编辑评估范围狭窄且存在偏差,仅关注特定任务,未能充分考察对领域内样本的影响。为此,我们提出ComprehendEdit,一个涵盖八个不同任务的综合性基准,来自多个数据集。我们引入两个新指标:知识泛化指数(KGI)和知识保留指数(KPI),用于评估编辑对领域内样本的影响,且无需依赖人工智能生成样本。基于该框架的洞察,我们建立分层上下文编辑(HICE)基线方法,采用两阶段策略,在各项指标间取得良好平衡。本研究提供了更全面的多模态知识编辑评估体系,揭示了该领域的独特挑战,并展示了性能提升的基线方法。我们的工作为未来研究开辟新视角,为开发更鲁棒、高效的多模态语言模型编辑技术奠定基础。ComprehendEdit基准及实现代码已开源:https://github.com/yaohui120/ComprehendEdit。

原文摘要 · Abstract (English)

Large multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current multimodal knowledge editing evaluations are limited in scope and potentially biased, focusing on narrow tasks and failing to assess the impact on in-domain samples. To address these issues, we introduce ComprehendEdit, a comprehensive benchmark comprising eight diverse tasks from multiple datasets. We propose two novel metrics: Knowledge Generalization Index (KGI) and Knowledge Preservation Index (KPI), which evaluate editing effects on in-domain samples without relying on AI-synthetic samples. Based on insights from our framework, we establish Hierarchical In-Context Editing (HICE), a baseline method employing a two-stage approach that balances performance across all metrics. This study provides a more comprehensive evaluation framework for multimodal knowledge editing, reveals unique challenges in this field, and offers a baseline method demonstrating improved performance. Our work opens new perspectives for future research and provides a foundation for developing more robust and effective editing techniques for MLLMs. The ComprehendEdit benchmark and implementation code are available at https://github.com/yaohui120/ComprehendEdit.

多模态知识编辑评估框架模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。