按病情严重程度分阶段训练,提升阿拉伯语医学文本生成效果
A Severity-Based Curriculum Learning Strategy for Arabic Medical Text Generation
- 按病情轻重分阶段训练,先学简单病例再攻复杂重症
- 在MAQA数据集上性能提升4%至7%,优于基线模型
- 适合医疗AI开发、多语言健康助手等场景使用
阿拉伯语医学文本生成日益重要,有助于用户以母语理解症状并获取健康建议。然而,现有方法通常假设所有训练样本重要性相同,忽略了临床严重程度的差异,这可能影响模型对复杂或高风险病例的捕捉能力。为此,本文提出一种基于严重程度的课程学习策略,将训练过程从较轻病情逐步过渡到危重病例。该方法依据严重程度对数据集进行分阶,细调时渐进引入更复杂的案例,使模型先掌握基础医学模式,再处理高难度情形。实验在包含症状描述与对应回答的阿拉伯语医学问答(MAQA)子集上进行,并通过本研究提出的规则方法标注了轻度、中度和重度三个严重等级。结果表明,引入严重程度感知的课程学习可显著提升所有测试模型表现,相较基线模型性能提升约4%至7%,相比传统微调方法提高3%至6%。
原文摘要 · Abstract (English)
Arabic medical text generation is increasingly needed to help users interpret symptoms and access general health guidance in their native language. Nevertheless, many existing methods assume uniform importance across training samples, overlooking differences in clinical severity. This simplification can hinder the model's ability to properly capture complex or high-risk cases. To overcome this issue, this work introduces a Severity-based Curriculum Learning Strategy for Arabic Medical Text Generation, where the training process is structured to move gradually from less severe to more critical medical conditions. The approach divides the dataset into ordered stages based on severity and incrementally exposes the model to more challenging cases during fine-tuning, allowing it to first learn basic medical patterns before addressing more complex scenarios. The proposed method is evaluated on a subset of the Medical Arabic Question Answering (MAQA) dataset, which includes Arabic medical questions describing symptoms alongside corresponding responses. In addition, the dataset is annotated with three severity levels (Mild, Moderate, and Critical) using a rule-based method developed in this study. The results demonstrate that incorporating severity-aware curriculum learning leads to consistent performance improvements across all tested models, with gains of around +4% to +7% over baseline models and +3% to +6% compared with conventional fine-tuning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。