为医疗视觉语言模型编辑设计了真实临床场景的评估基准
Evaluating and Understanding Model Editing for Medical Vision Language Models

- 构建面向医疗场景的多模态模型编辑评测集M3Bench
- 发现现有编辑方法在迁移性与局部性间存在根本权衡
- 适合关注医疗AI安全部署的研究者与开发者
模型编辑为在不重新训练的情况下快速修正医疗视觉语言模型(VLMs)的上线后错误提供了高效路径。然而,现有多模态模型编辑基准主要针对通用任务,无法反映真实临床场景的需求与变异性。为此,我们提出M3Bench——一个基于临床实际的多模态模型编辑评测基准,用于评估编辑在图像与文本变化、模态与协议转换、临床知识组合及时间演进等挑战下的可靠性、精确性与泛化能力。M3Bench包含16,276个问题,覆盖多种解剖结构、模态和专科领域,支持单次与连续编辑。通过对6种医学与通用VLMs上4种代表性编辑器的评估,发现无一方法在所有指标上表现优异:基于梯度的方法虽具强迁移性,但易引发灾难性局部性破坏;基于记忆的方法保持局部性,却缺乏组合泛化能力,且对主干网络敏感。我们进一步揭示这些失败源于VLM隐空间几何特性及不同编辑方式对其布局的影响。总体而言,M3Bench为多模态模型编辑提供了严格的临床压力测试,并为更安全的上线后适应提供实践指导。基准已开源于https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench。
原文摘要 · Abstract (English)
Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise, and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composition, and temporal progression. M3Bench contains 16,276 questions spanning diverse anatomy, modalities, and specialties, and supports both single and sequential edits. By evaluating 4 representative editors across 6 medical and general VLMs, we find that no method excels across all criteria. Gradient-based editors achieve strong transfer but suffer from catastrophic locality violations, whereas memory-based methods preserve locality but lack compositional generality and exhibit high backbone-dependent hyperparameter sensitivity. We further attribute these failures to the latent space geometry of VLMs and how different editing methods shift its landscape. Overall, M3Bench establishes a rigorous clinical stress test for multimodal model editing and offers actionable guidance for safer post-deployment adaptation. The benchmark is publicly available at https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。