arXiv:2509.08604cs.CLcs.AI2025-09被引 5

医学大模型普遍记忆训练数据,可能泄露患者信息且影响诊断可靠性。

Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications

  • 分析三类医学微调场景下的记忆现象,覆盖超1.3万份真实病历。
  • 记忆率显著高于通用领域,持续预训练后87%内容仍被保留。
  • 揭示记忆对临床应用的潜在风险,适合医疗AI安全研究者参考。

大型语言模型在医学领域展现出巨大潜力,通过持续预训练或微调医学数据提升专业准确性与安全性。然而,其对医学训练数据的记忆程度尚不明确。适度记忆可保留关键医学知识,但过度记忆可能导致敏感临床信息(如患者特异性细节)被无意复现,削弱模型泛化能力,增加误诊与不当建议风险。由于生成模型特性,记忆内容可能以自信但误导性输出呈现,阻碍临床采纳。本文系统研究医学大模型的记忆现象,评估其普遍性、特征、规模及下游影响。分析三种典型适应场景:(1)医学语料持续预训练;(2)标准医学基准微调;(3)真实临床数据微调,包含来自耶鲁纽黑文医疗系统的逾13,000份独立住院记录。结果表明,记忆现象在所有场景中普遍存在,且显著高于通用领域。持续预训练与微调阶段的记忆特征各异,且具有强持续性——高达87%的持续预训练记忆内容在后续任务微调后依然保留。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-training or fine-tuning on medical data to enhance domain-specific accuracy and safety. However, a key open question remains: to what extent do LLMs memorize medical training data. Memorization can be beneficial when it enables LLMs to retain valuable medical knowledge during domain adaptation. Yet, it also raises concerns. LLMs may inadvertently reproduce sensitive clinical content (e.g., patient-specific details), and excessive memorization may reduce model generalizability, increasing risks of misdiagnosis and making unwarranted recommendations. These risks are further amplified by the generative nature of LLMs, which can not only surface memorized content but also produce overconfident, misleading outputs that may hinder clinical adoption. In this work, we present a study on memorization of LLMs in medicine, assessing its prevalence (how frequently it occurs), characteristics (what is memorized), volume (how much content is memorized), and potential downstream impacts (how memorization may affect medical applications). We systematically analyze common adaptation scenarios: (1) continued pretraining on medical corpora, (2) fine-tuning on standard medical benchmarks, and (3) fine-tuning on real-world clinical data, including over 13,000 unique inpatient records from Yale New Haven Health System. The results demonstrate that memorization is prevalent across all adaptation scenarios and significantly higher than that reported in the general domain. Moreover, memorization has distinct characteristics during continued pre-training and fine-tuning, and it is persistent: up to 87% of content memorized during continued pre-training remains after fine-tuning on new medical tasks.

大模型记忆医学AI数据安全模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。