arXiv:2411.00897cs.CLcs.AI2024-11被引 5

用少量数据提升大模型中医能力,靠AI反馈强化学习。

Enhancing the Traditional Chinese Medicine Capabilities of Large Language Model through Reinforcement Learning from AI Feedback

  • 先用病例数据微调,再用AI反馈强化学习优化
  • 小样本下中医任务性能显著提升
  • 适合医疗AI、中医药数字化研究者

尽管大语言模型在理解用户意图方面表现良好,但在中医等专业领域仍因缺乏专业知识而表现受限。此外,高质量中医数据稀缺且难以获取,导致大模型在中医任务上效果不佳。本文提出一种仅需少量数据的框架,通过医学病例数据对大模型进行监督微调,使其初步具备中医处理能力;随后采用基于AI反馈的强化学习(RLAIF)进一步优化模型,使其与偏好数据对齐。消融实验表明性能提升来自监督微调和直接策略优化的共同作用。实验结果表明,使用少量数据训练的模型在代表性中医任务上实现了显著性能提升。

原文摘要 · Abstract (English)

Although large language models perform well in understanding and responding to user intent, their performance in specialized domains such as Traditional Chinese Medicine (TCM) remains limited due to lack of expertise. In addition, high-quality data related to TCM is scarce and difficult to obtain, making large language models ineffective in handling TCM tasks. In this work, we propose a framework to improve the performance of large language models for TCM tasks using only a small amount of data. First, we use medical case data for supervised fine-tuning of the large model, making it initially capable of performing TCM tasks. Subsequently, we further optimize the model's performance using reinforcement learning from AI feedback (RLAIF) to align it with the preference data. The ablation study also demonstrated the performance gain is attributed to both supervised fine-tuning and the direct policy optimization. The experimental results show that the model trained with a small amount of data achieves a significant performance improvement on a representative TCM task.

中医AI强化学习小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。