评估医疗大模型在真实威胁下的隐私泄露风险,发现病历信息易被还原。
Clinically Grounded Privacy Evaluation of Medical LMs

- 构建分级威胁评估框架,从公开信息到碎片病历逐级测试隐私泄露
- 模型对患者姓名、出生日期等信息记忆率达高,敏感诊断识别准确率高达0.91(堕胎)
- 发现36%的记忆内容为模板化文本,提示需谨慎解读记忆率
医疗语言模型可能记住并复现受保护的健康信息,但现有隐私评估多关注训练文本的恢复,而非真实威胁场景下的信息泄露。本文提出一种临床驱动的评估框架,从可公开推断的人口统计信息到泄露的病历片段,分层级评估模型的文本复现与语义泄露。在持续预训练于37.8万份临床记录的模型上测试发现,常规就诊元数据(如姓名、出生日期、就诊日期、医生姓名、机构位置)在患者整个时间线上引发高频率的原文记忆,且对敏感诊断的恢复表现优异(堕胎诊断的AUROC达0.91,HIV为0.82)。同时,精确匹配记忆可能夸大披露风险:36%的记忆词元来自模板化文档。研究揭示了长期临床数据训练带来的隐私风险,并提供了可复用的上下文化隐私评估框架。
原文摘要 · Abstract (English)
Medical language models (LMs) can memorize and reproduce protected health information, but privacy evaluations often focus on recovery of training text rather than disclosure under realistic threat models. We introduce a clinically grounded framework that evaluates leakage along a graded axis of adversarial access, ranging from publicly inferable demographics to leaked note fragments. At each tier, we measure verbatim memorization of patient-specific text and semantic leakage of sensitive diagnoses. Applying the framework to an LM continually pretrained on 378k clinical notes, we find that routine encounter metadata (i.e., name, date of birth, visit date, provider name, and practice location) elicits high rates of verbatim memorization across a patient's timeline and sensitive-diagnosis recovery (AUROC 0.91 for abortion, 0.82 for HIV). At the same time, exact-match memorization can overstate disclosure: 36% of memorized tokens reflect templated documentation. Our work highlights the risks of training on longitudinal clinical data and provides a practical, reusable framework for contextual privacy evaluation of medical LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。