arXiv:2605.03301cs.CLcs.AI2026-05

构建多样临床笔记数据集,训练可本地部署的高效去标识化小模型。

SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification

论文配图:SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
图 1 · 摘自论文原文
  • 基于多样性采样与人工校验构建1381篇真实临床笔记数据集
  • 蒸馏出在本地运行的微型模型,精确率达0.89,召回率达0.88
  • 适合需隐私保护、无法使用云端API的医疗机构使用

临床文本去标识化是电子健康记录二次利用的前提。现有公开基准如i2b2 2006和2014数据集已超十年,缺乏现代临床叙述的语义与人口多样性。大语言模型(LLMs)虽具备先进零样本提取能力,但因计算成本高及医院数据治理限制(禁止将受保护健康信息发送至云端接口),难以在企业级应用。我们提出SHIELD(用于学习与去标识化的合成人类标注替换条目),包含1,381篇笔记,共10,229个金标准隐私信息(PHI)片段,覆盖9类,通过跨人口与文档类型分层的集合覆盖多样性采样,并结合人工介入校验。评估四种LLM(两种专有、两种开源)以确立性能上限,随后展示教师-学生蒸馏框架可将这些能力迁移至本地部署的小型语言模型。最佳蒸馏模型在标准工作站上运行,微平均跨度级精确率为0.89,召回率为0.88;其分类层面召回率(0.90对0.81均值)略低于云上教师模型,但因其低成本与本地部署优势仍具竞争力。跨数据集评估表明,多样性训练模型在通用结构化隐私类别上泛化良好,而机构特异性实体在双向迁移中仍具挑战,建议对高通量半结构化笔记采用广覆盖模型与专用模型结合策略。我们公开发布SHIELD数据集与蒸馏后的DeBERTa v3模型,提供完全可在机构防火墙内部署的准确、经济的去标识化方案。

原文摘要 · Abstract (English)

De-identification of clinical text is a prerequisite for the secondary use of electronic health records. Existing public benchmarks such as the i2b2 2006 and 2014 corpora are over a decade old and lack the semantic and demographic diversity of modern clinical narratives. Large Language Models (LLMs) reach state-of-the-art zero-shot extraction, but their use at enterprise scale is limited by computational cost and by hospital data governance that restricts sending Protected Health Information (PHI) to cloud APIs. We introduce SHIELD (Synthetic Human-annotated Identifier-replaced Entries for Learning and De-identification), a diverse clinical note dataset of 1,381 notes with 10,229 gold-standard PHI spans across 9 categories, built with set-cover diversity sampling across demographic and document-type strata and human-in-the-loop adjudication. We evaluate four LLMs (two proprietary, two open-weight) to establish a performance ceiling on SHIELD, then show that a teacher-student distillation framework transfers these capabilities into locally deployable Small Language Models. Our best distilled model reaches micro-averaged span-level precision of 0.89 and recall of 0.88 while running on standard workstation hardware. It trails its cloud teacher on per-category recall (0.90 vs. 0.81 macro-averaged) but remains competitive given its lower cost and on-premise deployability. Cross-dataset evaluation shows that diversity-trained models generalize well on universal structured PHI categories, while institution-specific entities remain hard to transfer in both directions, which suggests pairing broad-coverage models with specialized models for high-volume, semi-structured note types. We publicly release the SHIELD dataset and the distilled DeBERTa v3 model to provide an accurate, cost-effective de-identification pipeline deployable entirely behind institutional firewalls.

医疗AI去标识化小模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。