arXiv:2506.03490cs.CL2025-06Conference of the …被引 4

测试医学大模型知识编辑效果,发现现有方法只记表面信息。

Beyond Memorization: A Rigorous Evaluation Framework for Medical Knowledge Editing

  • 用医学知识编辑基准测试评估不同编辑方式
  • 现有方法仅表面记忆,无法应对新场景
  • 提出自生成推理编辑法,提升泛化能力

知识编辑(KE)近年来成为在不重新训练大型语言模型(LLMs)的情况下更新特定事实的有前景方法。尽管在通用领域表现良好,其在复杂医学领域的适用性仍待探索。医学知识编辑尤为困难,要求模型内化知识并泛化至未见情境以实现有效且可解释的决策。本文提出新型框架MedEditBench,用于严格评估现有KE方法在医学领域的有效性。该框架包含新的医学知识编辑基准及三种不同的编辑范式,旨在评估不同知识源的影响。研究发现,当前KE方法仅导致注入信息的浅层记忆,无法泛化至新场景。为此,我们提出自生成推理编辑(SGR-Edit),利用模型生成的推理作为编辑目标,揭示底层推理过程,并显著优于现有方法。此外,我们深入探讨了医学知识在LLMs中的定位及序列编辑对动态知识演化的影响,为实际医疗应用中实施KE提供指导。

原文摘要 · Abstract (English)

Recently, knowledge editing (KE) has emerged as a promising approach to update specific facts in Large Language Models (LLMs) without the need for full retraining. Despite the effectiveness in general-domain benchmarks, their applicability to complex medical domain remains largely unexplored. Medical knowledge editing is particularly challenging, as it requires LLMs to internalize the knowledge and generalize to unseen scenarios for effective and interpretable decision-making. In this work, we propose a novel framework called MedEditBench to rigorously evaluate the effectiveness of existing KE methods in the medical domain. In MedEditBench, we introduce a new medical knowledge editing benchmark as well as three different knowledge editing paradigms, which are designed to assess the impact of different knowledge sources for editing. Our findings indicate that current KE methods result in only superficial memorization of the injected information, failing to generalize to new scenarios. To overcome this limitation, we present Self-Generated Rationale Editing (SGR-Edit), which utilizes model-derived rationales as the target knowledge for editing, thereby uncovering the underlying reasoning process and demonstrating significant improvements over existing KE approaches. Additionally, we offer deeper insights into medical knowledge editing, including the localization of medical knowledge in LLMs and the impact of sequential editing on evolving knowledge. This could provide practical guidance for implementing KE methods in real-world medical applications.

知识编辑医学AI大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。