arXiv:2501.02869cs.CLcs.AI2025-01被引 4

用高效偏好对齐提升医疗大模型对话能力

IIMedGPT: Promoting Large Language Model Capabilities of Medical Tasks by Efficient Human Preference Alignment

  • 构建六类真实医疗任务指令数据集CMedINS,用于医学大模型微调
  • 采用DPO方法实现高效偏好对齐,模型在医疗对话中表现优于现有模型
  • 适合医疗AI研究者、临床辅助系统开发者参考

近期基于大规模通用语料预训练的大语言模型在回应人类查询方面取得突破,但面临数据不足和无法对齐用户指令的挑战。为此,我们构建了一个包含六类真实医疗任务指令的医学指令数据集CMedINS,结合其他数据有效微调大语言模型。随后,我们推出了医学模型IIMedGPT,采用高效的直接偏好优化(Direct Preference Optimization, DPO)方法。实验结果表明,该模型在医疗对话任务中优于现有医学模型。数据集、代码及模型检查点将在论文被接受后公开。

原文摘要 · Abstract (English)

Recent researches of large language models(LLM), which is pre-trained on massive general-purpose corpora, have achieved breakthroughs in responding human queries. However, these methods face challenges including limited data insufficiency to support extensive pre-training and can not align responses with users' instructions. To address these issues, we introduce a medical instruction dataset, CMedINS, containing six medical instructions derived from actual medical tasks, which effectively fine-tunes LLM in conjunction with other data. Subsequently, We launch our medical model, IIMedGPT, employing an efficient preference alignment method, Direct preference Optimization(DPO). The results show that our final model outperforms existing medical models in medical dialogue.Datsets, Code and model checkpoints will be released upon acceptance.

医疗大模型偏好对齐DPO指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。