arXiv:2608.17051cs.CLcs.AI2026-08

用特定机构提示词让大模型找回被漏掉的医疗隐私信息

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

论文配图:Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss
图 1 · 摘自论文原文
  • 用上下文学习设计机构定制提示,精准识别本地化医疗敏感信息
  • 召回率达98.1%,比现有系统提升显著,且避免过度删除临床内容
  • 适合需要高精度隐私保护的医疗机构自研标注体系

电子病历二次使用需去标识化,但现有系统会遗漏医院缩写、楼栋名等具有机构特异性的受保护健康信息(PHI),其归属由本地决定。我们测试大语言模型(LLMs)结合上下文学习(ICL)是否能弥补这一缺口并调控精确率-召回率权衡。在德克萨斯儿童医院100份儿科肿瘤病历(5,322个PHI片段)上,对比了8个LLM与两个专用系统(Stanford TiDE、OpenMed PII)及两个基于规则的基线。每个LLM采用三类逐步细化的提示:(1)符合HIPAA的基准提示,(2)加入其遗漏的机构类别,(3)在第二步基础上增加防止过度脱敏的指令。再比较14种多智能体与集成配置与最优单提示表现,以召回率为首要安全指标。结果表明,LLMs超越专用系统(最佳F1=0.918±0.001 vs. TiDE 0.779),优势集中在上下文类信息。命名遗漏类别后恢复了79%(48/61)的缺失项,抑制过度脱敏显著提升精确率。无智能体架构优于校准后的单次提示(F1 0.906–0.907),但LLM输出揭示414个候选标注空白;重新标注确认227个新增PHI,最终提示达到召回率0.981(F1=0.907±0.002)。经校准的ICL可在一次调用中同时解决机构性隐私遗漏与精确率-召回率权衡问题。尽管运行成本高于传统方法,但可实现对参考标准的审计。LLMs是专用去标识化系统的合法且灵活替代方案,机构特定提示开发应为首要适配策略。

原文摘要 · Abstract (English)

Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.

医疗隐私大模型去标识化提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。