用本地微调大模型实现病历隐私信息精准标注,兼顾安全与效率。
Large Language Model Empowered Privacy-Protected Framework for PHI Annotation in Clinical Notes
- 本地微调合成病历数据的LLM,避免云端隐私泄露风险。
- 在真实病历上达到高精度隐私信息识别,准确率优于传统方法。
- 适合医疗数据隐私保护场景,尤其关注数据安全的研究者使用。
医学数据中的隐私信息去标识化是降低保密泄露风险的关键步骤,尤其当患者个人细节未被充分移除时。尽管已有基于规则和学习的方法,但普遍存在泛化能力弱、需大量标注数据的问题。近年来,大语言模型(LLMs)凭借其出色的语义理解能力展现出解决该问题的潜力。然而,使用商业LLM API存在隐私风险,而本地部署开源LLM又带来高昂计算成本。本文提出一种基于大语言模型的隐私保护病历标注框架LPPA,专用于英文临床笔记。通过在本地对LLM进行合成病历数据的微调,LPPA实现了强隐私保护与高精度的隐私信息标注。大量实验表明,LPPA能有效精准地去标识私密信息,提供一种可扩展且高效的患者隐私保护方案。
原文摘要 · Abstract (English)
The de-identification of private information in medical data is a crucial process to mitigate the risk of confidentiality breaches, particularly when patient personal details are not adequately removed before the release of medical records. Although rule-based and learning-based methods have been proposed, they often struggle with limited generalizability and require substantial amounts of annotated data for effective performance. Recent advancements in large language models (LLMs) have shown significant promise in addressing these issues due to their superior language comprehension capabilities. However, LLMs present challenges, including potential privacy risks when using commercial LLM APIs and high computational costs for deploying open-source LLMs locally. In this work, we introduce LPPA, an LLM-empowered Privacy-Protected PHI Annotation framework for clinical notes, targeting the English language. By fine-tuning LLMs locally with synthetic notes, LPPA ensures strong privacy protection and high PHI annotation accuracy. Extensive experiments demonstrate LPPA's effectiveness in accurately de-identifying private information, offering a scalable and efficient solution for enhancing patient privacy protection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。