arXiv:2505.03025cs.CLcs.AI2025-05中稿 · LREC 2026 https://…被引 2

为临床对话合成数据制定分类体系,助力医疗NLP研究

A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts

  • 提出临床对话合成数据的新型分类框架
  • 梳理合成数据生成与评估方法,覆盖多类医疗场景
  • 帮助研究者选择合适数据,提升模型泛化能力

合成数据在语言学领域和自然语言处理任务中广泛应用,尤其在真实数据稀缺或缺失的情况下。医疗领域即为典型例子,由于隐私保护、匿名化和数据治理等长期挑战,促使合成数据集不断涌现。其中,临床对话数据尤为敏感且难以采集,常通过合成方式生成。尽管已有研究证明其在某些情境下具备足够有效性,但缺乏理论指导以明确如何最佳使用及推广至新应用。本文系统综述了医疗对话任务中合成数据的构建、评估与应用现状,并提出一种新颖的合成数据分类体系,用于区分合成类型与程度,从而促进不同数据集之间的比较与评估。

原文摘要 · Abstract (English)

Synthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of clinical (healthcare) contexts, where there exist significant and long-standing challenges (e.g., privacy, anonymization, and data governance) which have led to the development of an increasing number of synthetic datasets. One increasingly important category of clinical dataset is that of clinical dialogues which are especially sensitive and difficult to collect, and as such are commonly synthesized. While such synthetic datasets have been shown to be sufficient in some situations, little theory exists to inform how they may be best used and generalized to new applications. In this paper, we provide an overview of how synthetic datasets are created, evaluated and being used for dialogue related tasks in the medical domain. Additionally, we propose a novel typology for use in classifying types and degrees of data synthesis, to facilitate comparison and evaluation.

合成数据医疗NLP对话系统数据分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。