arXiv:2603.14265cs.CLcs.MA2026-03中稿 · EMNLP被引 1

首个医疗问答隐私-效用平衡评测基准,揭示大模型的隐私泄露风险。

MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-Ended Question Answering

  • 构建多智能体人类协作管道生成真实医疗隐私场景
  • 九个模型测试显示隐私与效用存在普遍权衡,最高泄露率达72.8%
  • 适合医疗AI安全评估、隐私保护算法研发者参考

检索增强生成使大模型能基于临床证据生成回答,但连接外部数据库可能导致上下文泄露,即通过独特的医疗组合实现患者再识别。现有医疗评测侧重准确率,忽视此风险。为此,我们提出MedPriv-Bench,首个联合评估医疗开放问答中隐私保护与临床效用的基准。框架采用多智能体、人机协同流程,合成敏感医疗背景与临床相关问题以制造真实隐私压力。我们建立自动化评估协议,使用微调后的RoBERTa-NLI模型,对人类标注的实例达到75.3%的F1分数、90.7%的灵敏度,平均推理时间0.056秒/样本。在九个大模型和三种隐私保护方法中,观察到普遍的隐私-效用权衡:相比未防护的Med42-v2-8B(效用3.87/5;泄露72.8%),监督微调将效用提升至4.25,泄露降至38.9%;而本地差分隐私将泄露降至20.5%,但效用下降至3.03。结果表明,需建立领域专用基准来验证医疗AI在隐私敏感场景下的可靠性。

原文摘要 · Abstract (English)

Recent advances in Retrieval-Augmented Generation enable LLMs to ground outputs in clinical evidence, but connections to external databases create the risk of contextual leakage, where unique combinations of medical details enable patient re-identification without explicit identifiers. Existing healthcare benchmarks emphasize accuracy while overlooking this risk. To fill this gap, we present MedPriv-Bench, the first benchmark for jointly evaluating privacy preservation and clinical utility in medical open-ended question answering. Our framework utilizes a multi-agent, human-in-the-loop pipeline to synthesize sensitive medical contexts and clinically relevant queries that create realistic privacy pressure. We also establish an automated evaluation protocol using a fine-tuned RoBERTa-NLI model, which achieved an instance-level F1 score of 75.3%, sensitivity of 90.7%, and an average inference time of 0.056 s per sample against human annotations. Across nine LLMs and three privacy-preserving methods, we observed a pervasive privacy-utility trade-off. Relative to unprotected Med42-v2-8B (utility 3.87/5; leakage 72.8%), supervised fine-tuning improved utility to 4.25 and reduced leakage to 38.9%, whereas local differential privacy reduced leakage to 20.5% but lowered utility to 3.03. These results demonstrate the need for domain-specific benchmarks to validate medical AI systems in privacy-sensitive settings.

医疗AI隐私保护大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。