arXiv:2502.14677cs.CLcs.AI2025-02ACL被引 3

用大模型生成临床文本,实现隐私保护下的命名实体识别训练。

Data-Constrained Synthesis of Training Data for De-Identification

  • 用领域适配的LLM生成带敏感信息标签的合成临床文本。
  • 合成数据训练的NER模型性能仅下降约5%-8%。
  • 适合缺乏标注数据的医疗等敏感领域研究者使用。

许多敏感领域(如临床)因隐私风险而缺乏可用数据集。大语言模型(LLMs)的生成能力使合成数据成为可行方案。本研究将LLM适配至临床领域,生成带有个人身份信息(PII)标签的合成临床文本,并利用高性能编码器型NER模型进行机器标注。这些合成语料用于训练合成NER模型。结果表明,使用合成语料训练的NER模型预测性能仅略有下降。通过瑞典语和西班牙语数据的系统消融实验,发现小规模数据即可完成LLM的领域适配,而该流程的有效性几乎完全依赖于基于原始数据训练的机器标注NER模型性能。

原文摘要 · Abstract (English)

Many sensitive domains -- such as the clinical domain -- lack widely available datasets due to privacy risks. The increasing generative capabilities of large language models (LLMs) have made synthetic datasets a viable path forward. In this study, we domain-adapt LLMs to the clinical domain and generate synthetic clinical texts that are machine-annotated with tags for personally identifiable information using capable encoder-based NER models. The synthetic corpora are then used to train synthetic NER models. The results show that training NER models using synthetic corpora incurs only a small drop in predictive performance. The limits of this process are investigated in a systematic ablation study -- using both Swedish and Spanish data. Our analysis shows that smaller datasets can be sufficient for domain-adapting LLMs for data synthesis. Instead, the effectiveness of this process is almost entirely contingent on the performance of the machine-annotating NER models trained using the original data.

数据合成隐私保护NER

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。