arXiv:2412.04573cs.CL2024-12中稿 · ML4H 2024 Findings被引 6

用大模型生成更难的临床问答数据,提升医疗AI训练效果

Give me Some Hard Questions: Synthetic Data Generation for Clinical QA

  • 通过规避上下文重叠和预设模板引导,生成更具挑战性的临床问题
  • 在两个数据集上使微调性能显著优于基线方法
  • 适合医疗AI研究者构建高质量合成数据集

临床问答系统可帮助医生快速从电子健康记录中获取患者信息。但训练此类系统需大量标注数据,而临床数据因专业性强且涉及隐私,难以获取。本文探索在零样本设置下利用大语言模型生成临床问答数据。发现简单提示常产生过于简单的问答对,无法反映真实临床复杂性。为此提出两种提示策略:1)指令模型生成与输入上下文无重叠的问题;2)采用预定义结构化模板总结病历以引导问题生成。在两个临床问答数据集上的实验表明,该方法生成的问题更具挑战性,显著提升微调性能。对比合成数据与真实数据发现,两者在训练效果上的差距源于合成答案的质量差异。

原文摘要 · Abstract (English)

Clinical Question Answering (QA) systems enable doctors to quickly access patient information from electronic health records (EHRs). However, training these systems requires significant annotated data, which is limited due to the expertise needed and the privacy concerns associated with clinical data. This paper explores generating Clinical QA data using large language models (LLMs) in a zero-shot setting. We find that naive prompting often results in easy questions that do not reflect the complexity of clinical scenarios. To address this, we propose two prompting strategies: 1) instructing the model to generate questions that do not overlap with the input context, and 2) summarizing the input record using a predefined schema to scaffold question generation. Experiments on two Clinical QA datasets demonstrate that our method generates more challenging questions, significantly improving fine-tuning performance over baselines. We compare synthetic and gold data and find a gap between their training efficacy resulting from the quality of synthetically generated answers.

临床QA合成数据大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。