用真实病历嵌入引导生成更像临床文本的合成数据
Embedding-Driven Diversity Sampling to Improve Few-Shot Synthetic Data Generation
- 从少量真实病历中采样多样性嵌入,指导大模型生成
- 合成数据使分类性能提升,降低40%数据需求
- 适合医疗文本生成与少样本学习研究者
临床文本准确分类通常需微调预训练语言模型,但依赖高质量数据和专家标注,成本高昂。合成数据生成可作为替代方案,但预训练模型难以捕捉临床笔记的句法多样性。本文提出一种基于嵌入的多样性采样方法,利用少量真实临床笔记的嵌入信息,引导大语言模型在少样本提示下生成更贴近真实临床语法的合成文本。在CheXpert数据集上进行分类任务评估,对比随机少样本与零样本方法,采用余弦相似度和图灵测试验证,结果表明该方法生成的合成笔记更接近真实临床文本。其管道使达到0.85 AUC阈值所需数据量减少40%(AUROC)和30%(AUPRC),且合成数据增强模型后,AUROC提升57%,AUPRC提升68%。此外,合成数据效能达真实数据的0.9倍,价值提升60%。
原文摘要 · Abstract (English)
Accurate classification of clinical text often requires fine-tuning pre-trained language models, a process that is costly and time-consuming due to the need for high-quality data and expert annotators. Synthetic data generation offers an alternative, though pre-trained models may not capture the syntactic diversity of clinical notes. We propose an embedding-driven approach that uses diversity sampling from a small set of real clinical notes to guide large language models in few-shot prompting, generating synthetic text that better reflects clinical syntax. We evaluated this method using the CheXpert dataset on a classification task, comparing it to random few-shot and zero-shot approaches. Using cosine similarity and a Turing test, our approach produced synthetic notes that more closely align with real clinical text. Our pipeline reduced the data needed to reach the 0.85 AUC cutoff by 40% for AUROC and 30% for AUPRC, while augmenting models with synthetic data improved AUROC by 57% and AUPRC by 68%. Additionally, our synthetic data was 0.9 times as effective as real data, a 60% improvement in value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。