arXiv:2606.05510cs.AIcs.CL2026-06

按病情严重程度分阶段训练,提升医疗问答生成质量

Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation

  • 按轻中重三类病情分阶段训练多模型,逐步积累医学知识
  • 在MAQA数据集上,生成结果的BERTScore达90.30%(微调后)
  • 适合需要高准确性和情境适应性的医疗AI系统使用

远程医疗系统在提供可及且及时的医疗信息方面日益重要。现有大语言模型在不同病情严重程度下难以保持一致且合适的医学回应。为此,我们提出一种基于严重程度感知的多模型框架,结合课程学习与基于相关性的响应选择策略。该框架采用三阶段课程学习:各模型依次在轻度、中度和重度病例上训练,逐步获取领域知识。使用五个独立训练的大语言模型,在推理时生成候选回复,选取BERTScore最高的作为最终输出。框架在包含标注问答对的MAQA数据集上进行训练与评估。实验结果表明,该方法在BERTScore上优于基线与微调模型,在基线设置下达到86.71%,微调后提升至90.30%。结果验证了课程学习与多模型响应选择结合在提升医疗文本生成质量与相关性方面的有效性。

原文摘要 · Abstract (English)

Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language models often struggle to provide consistent and contextually appropriate medical responses across varying levels of case severity. This limitation highlights the need for models that can effectively adapt to the progressive complexity in medical queries. To address this challenge, we introduce a severity-aware multi-model framework that integrates curriculum training strategy with relevance-based response selection. The proposed framework employs a three-stage curriculum learning strategy, where each model is trained sequentially on mild, moderate, and critical cases to progressively acquire domain knowledge. The approach uses five large language models, each trained independently under the same curriculum. During inference, all models generate candidate responses, and the response with highest BERTScore is selected as the final output. The framework is trained and evaluated on the MAQA dataset, which provides annotated medical question-answer pairs. Experimental results evaluated using BERTScore demonstrate that the proposed method achieves superior performance compared to both baseline and fine-tuned models, attaining 86.71% in the baseline setting and 90.30% after fine-tuning. These results highlight the effectiveness of combining curriculum learning with multi-model response selection in improving response quality and relevance in medical text generation.

医疗问答课程学习多模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。