arXiv:2508.05083cs.AI2025-08AAAI被引 2

首个医学多模态知识编辑基准,评估模型更新医疗知识的能力

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

  • 构建包含反事实修正等四类任务的医学多模态知识编辑测试集
  • 实测主流模型在连续编辑中效果显著下降,暴露现有方法局限性
  • 适合医疗AI研究者与多模态模型开发者参考

多模态大语言模型(MLLM)在医疗AI领域取得显著进展,能融合视觉与文本信息。然而,随着医学知识不断更新,亟需在不从头训练的前提下高效修正模型中的过时或错误信息。尽管文本知识编辑已受广泛研究,但涉及图像与文本模态的医学多模态知识编辑仍缺乏系统性评估基准。为此,我们提出MedMKEB,首个综合性医学多模态知识编辑基准,用于评估知识编辑在可靠性、泛化性、局部性、可迁移性与鲁棒性方面的表现。该基准基于高质量医学视觉问答数据集,设计了反事实修正、语义泛化、知识迁移及对抗鲁棒性等编辑任务,并引入专家人工验证确保准确性。对先进通用与医学MLLM进行单次与序列编辑实验发现,现有基于知识的编辑方法在医学场景中存在明显不足,凸显发展专用编辑策略的必要性。MedMKEB将作为标准基准,推动可信、高效的医学知识编辑算法发展。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have significantly improved medical AI, enabling it to unify the understanding of visual and textual information. However, as medical knowledge continues to evolve, it is critical to allow these models to efficiently update outdated or incorrect information without retraining from scratch. Although textual knowledge editing has been widely studied, there is still a lack of systematic benchmarks for multimodal medical knowledge editing involving image and text modalities. To fill this gap, we present MedMKEB, the first comprehensive benchmark designed to evaluate the reliability, generality, locality, portability, and robustness of knowledge editing in medical multimodal large language models. MedMKEB is built on a high-quality medical visual question-answering dataset and enriched with carefully constructed editing tasks, including counterfactual correction, semantic generalization, knowledge transfer, and adversarial robustness. We incorporate human expert validation to ensure the accuracy and reliability of the benchmark. Extensive single editing and sequential editing experiments on state-of-the-art general and medical MLLMs demonstrate the limitations of existing knowledge-based editing approaches in medicine, highlighting the need to develop specialized editing strategies. MedMKEB will serve as a standard benchmark to promote the development of trustworthy and efficient medical knowledge editing algorithms.

医学AI多模态知识编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。