arXiv:2508.07022cs.AIcs.CL2025-08

首个面向医学多模态问答的知识编辑评估基准

MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

  • 构建多模态医学问答场景下的知识编辑评测框架
  • 发现现有方法在复杂临床流程中泛化能力差
  • 适合医疗AI安全与可解释性研究者参考

知识编辑(KE)为在不重新训练的情况下更新大语言模型的事实知识提供了可扩展的方法。尽管先前研究在通用领域和医学问答任务中已证明其有效性,但对多模态医学场景下的知识编辑关注较少。与纯文本环境不同,医学知识编辑需将更新后的知识与视觉推理结合,以支持安全且可解释的临床决策。为此,我们提出MultiMedEdit,首个针对临床多模态任务中知识编辑评估的基准。该框架涵盖理解与推理两类任务,定义了三维度量体系(可靠性、泛化性、局部性),并支持跨范式比较,涵盖通用模型与领域特定模型。我们在单次编辑与终身编辑设置下进行了广泛实验。结果表明,当前方法在泛化性和长尾推理方面表现不佳,尤其在复杂临床工作流中。我们还进行了效率分析(如编辑延迟、内存占用),揭示了不同知识编辑范式在实际部署中的权衡。总体而言,MultiMedEdit不仅揭示了现有方法的局限性,也为未来开发临床稳健的知识编辑技术奠定了基础。

原文摘要 · Abstract (English)

Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical QA tasks, little attention has been paid to KE in multimodal medical scenarios. Unlike text-only settings, medical KE demands integrating updated knowledge with visual reasoning to support safe and interpretable clinical decisions. To address this gap, we propose MultiMedEdit, the first benchmark tailored to evaluating KE in clinical multimodal tasks. Our framework spans both understanding and reasoning task types, defines a three-dimensional metric suite (reliability, generality, and locality), and supports cross-paradigm comparisons across general and domain-specific models. We conduct extensive experiments under single-editing and lifelong-editing settings. Results suggest that current methods struggle with generalization and long-tail reasoning, particularly in complex clinical workflows. We further present an efficiency analysis (e.g., edit latency, memory footprint), revealing practical trade-offs in real-world deployment across KE paradigms. Overall, MultiMedEdit not only reveals the limitations of current approaches but also provides a solid foundation for developing clinically robust knowledge editing techniques in the future.

知识编辑医学AI多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。