在本地运行小模型,安全高效地识别医疗缩写,防止误诊风险。
PLACID: Privacy-preserving Large language models for Acronym Clinical Inference and Disambiguation
- 分层处理:先用通用模型检出缩写,再由医学专用模型精准展开
- 本地部署2B-10B参数模型,缩写展开准确率达81%
- 兼顾隐私与效果,适合医院等敏感场景使用
大型语言模型在诸多领域展现变革潜力,但医疗应用受限于严格的数据隐私要求。临床文本中充斥着易混淆的缩写,误读可能导致致命用药错误。尽管云端大模型在缩写消歧上表现优异,但将受保护健康信息传至外部服务器违反隐私规范。为此,本研究首次评估在设备端完全部署的小参数模型以保障隐私。提出一种隐私保护级联流程:先用通用本地模型检测临床缩写,再将其路由至领域特定生物医学模型进行上下文相关展开。结果显示,通用指令跟随模型检测准确率高达~0.988,但展开能力下降至~0.655;而级联方案通过领域模型将展开准确率提升至~0.81。本研究证明,基于2B-10B参数的本地化模型可实现高保真临床缩写消歧支持。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer transformative solutions across many domains, but healthcare integration is hindered by strict data privacy constraints. Clinical narratives are dense with ambiguous acronyms, misinterpretation these abbreviations can precipitate severe outcomes like life-threatening medication errors. While cloud-dependent LLMs excel at Acronym Disambiguation, transmitting Protected Health Information to external servers violates privacy frameworks. To bridge this gap, this study pioneers the evaluation of small-parameter models deployed entirely on-device to ensure privacy preservation. We introduce a privacy-preserving cascaded pipeline leveraging general-purpose local models to detect clinical acronyms, routing them to domain-specific biomedical models for context-relevant expansions. Results reveal that while general instruction-following models achieve high detection accuracy (~0.988), their expansion capabilities plummet (~0.655). Our cascaded approach utilizes domain-specific medical models to increase expansion accuracy to (~0.81). This novel work demonstrates that privacy-preserving, on-device (2B-10B) models deliver high-fidelity clinical acronym disambiguation support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。