arXiv:2509.21530cs.LG2025-09中稿 · ICML

用专家指导的问答协作,让大模型生成更安全的医疗文本数据。

Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

  • 通过问答协作框架引入专家知识,引导大模型生成医疗文本。
  • 生成数据在关键医学信息保留上显著优于现有方法,幻觉减少。
  • 适合医疗领域需高安全性的文本增强任务,如临床预测建模。

数据增强是通过合成样本丰富训练数据以提升模型鲁棒性和泛化能力的常用策略。尽管大语言模型(LLMs)在生成方面表现出强大能力,但在医疗等高风险领域应用时,存在生成临床错误或误导性信息的风险。本文提出一种基于查询的模型协作框架,融入专家级领域知识,指导增强过程以保留关键医学信息。相比现有基于LLM和传统的方法,本方法生成的数据在词元和概念层面均显著提升了关键医学信息的保留率,减少了幻觉现象。下游临床预测任务的实验表明,该方法持续优于现有增强方法。这一轻量级协作框架有效弥合了大模型增强潜力与专业领域安全需求之间的差距。

原文摘要 · Abstract (English)

Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Compared to existing LLM-based and traditional augmentation methods, our generated data significantly improves preservation of critical medical information and reduces hallucinations at both the token and concept levels. Experiments on downstream clinical prediction tasks demonstrate consistent performance gains over existing augmentation methods. This lightweight collaborative framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.

医疗文本数据增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。