arXiv:2510.18674cs.CRcs.AI2025-10中稿 · the 1st IEEE Works…被引 1

发现临床大模型存在患者数据泄露风险,可能被攻击者推断病历是否参与训练。

Exploring Membership Inference Vulnerabilities in Clinical Large Language Models

  • 用问答模型测试隐私漏洞,对比传统攻击与临床场景更真实的改写攻击
  • 初步发现存在可测量的成员推理泄露,但程度有限
  • 适合关注医疗AI隐私安全的研究者和开发者参考

随着大型语言模型(LLMs)在临床决策支持、病历记录和患者信息系统的深度应用,其隐私与可信性已成为医疗领域的紧迫挑战。在敏感电子健康记录(EHR)数据上微调LLMs虽提升领域适配性,但也增加通过模型行为泄露患者信息的风险。本文开展一项探索性实证研究,评估临床LLM中的成员推理脆弱性,重点考察攻击者能否推断特定患者记录是否用于模型训练。采用先进的临床问答模型Llemr,我们评估了基于损失的经典攻击方法,以及更贴近临床对抗场景的改写扰动策略。初步结果表明存在有限但可测量的成员推理泄露,说明当前临床LLM具有部分抗性,但仍面临细微隐私风险,可能削弱临床AI的可信度。该发现推动发展上下文感知、领域特定的隐私评估与防御机制,如差分隐私微调和改写感知训练,以增强医疗AI系统安全性与可信性。

原文摘要 · Abstract (English)

As large language models (LLMs) become progressively more embedded in clinical decision-support, documentation, and patient-information systems, ensuring their privacy and trustworthiness has emerged as an imperative challenge for the healthcare sector. Fine-tuning LLMs on sensitive electronic health record (EHR) data improves domain alignment but also raises the risk of exposing patient information through model behaviors. In this work-in-progress, we present an exploratory empirical study on membership inference vulnerabilities in clinical LLMs, focusing on whether adversaries can infer if specific patient records were used during model training. Using a state-of-the-art clinical question-answering model, Llemr, we evaluate both canonical loss-based attacks and a domain-motivated paraphrasing-based perturbation strategy that more realistically reflects clinical adversarial conditions. Our preliminary findings reveal limited but measurable membership leakage, suggesting that current clinical LLMs provide partial resistance yet remain susceptible to subtle privacy risks that could undermine trust in clinical AI adoption. These results motivate continued development of context-aware, domain-specific privacy evaluations and defenses such as differential privacy fine-tuning and paraphrase-aware training, to strengthen the security and trustworthiness of healthcare AI systems.

医疗AI隐私安全成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。