arXiv:2507.18451cs.CLcs.AI2025-07综述被引 8

用合成医学文本解决数据稀缺与隐私问题,提升临床NLP效果。

Generation of Synthetic Clinical Text: A Systematic Review

  • 采用Transformer(尤其GPT)生成合成医疗自由文本,用于数据增强与隐私保护。
  • 合成文本在相似性、结构和实用性上表现良好,能有效缓解数据稀疏问题。
  • 适合需要安全数据集的研究者,但需人工审核避免敏感信息泄露。

生成临床合成文本是应对临床NLP中数据稀疏与隐私问题的有效方案。本文系统综述了2018年以来生成合成医学自由文本的研究,针对生成目的、技术方法与评估方式三个核心问题进行量化分析。从PubMed、ScienceDirect、Web of Science、Scopus、IEEE、Google Scholar和arXiv共检索1,398篇文献,筛选出94篇相关研究。主要生成目的包括文本增广、辅助写作、语料库构建、隐私保护、标注及实用性。主流技术为Transformer架构,尤其是GPT系列模型。评估主要涵盖相似性、隐私性、结构与实用性四个维度,其中实用性评价最常见。尽管合成文本在下游NLP任务中表现出中等程度的真实性,但已证明可作为真实文档的补充,显著提升模型准确率并缓解数据不足问题。然而,隐私风险仍存,需更多人工核查以排除敏感信息。未来合成文本技术将大幅加速临床工作流与系统开发,减少繁琐的数据合规流程。

原文摘要 · Abstract (English)

Generating clinical synthetic text represents an effective solution for common clinical NLP issues like sparsity and privacy. This paper aims to conduct a systematic review on generating synthetic medical free-text by formulating quantitative analysis to three research questions concerning (i) the purpose of generation, (ii) the techniques, and (iii) the evaluation methods. We searched PubMed, ScienceDirect, Web of Science, Scopus, IEEE, Google Scholar, and arXiv databases for publications associated with generating synthetic medical unstructured free-text. We have identified 94 relevant articles out of 1,398 collected ones. A great deal of attention has been given to the generation of synthetic medical text from 2018 onwards, where the main purpose of such a generation is towards text augmentation, assistive writing, corpus building, privacy-preserving, annotation, and usefulness. Transformer architectures were the main predominant technique used to generate the text, especially the GPTs. On the other hand, there were four main aspects of evaluation, including similarity, privacy, structure, and utility, where utility was the most frequent method used to assess the generated synthetic medical text. Although the generated synthetic medical text demonstrated a moderate possibility to act as real medical documents in different downstream NLP tasks, it has proven to be a great asset as augmented, complementary to the real documents, towards improving the accuracy and overcoming sparsity/undersampling issues. Yet, privacy is still a major issue behind generating synthetic medical text, where more human assessments are needed to check for the existence of any sensitive information. Despite that, advances in generating synthetic medical text will considerably accelerate the adoption of workflows and pipeline development, discarding the time-consuming legalities of data transfer.

合成数据临床NLP隐私保护文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。