arXiv:2510.17402cs.CLcs.AI2025-10被引 1

用新强化学习方法训练首个中医专用大模型,提升推理准确性。

Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine

  • 采用组内相对策略优化,通过对比选择改善回答质量
  • 在多个推理指标上超越GPT-4、Claude 3等主流模型
  • 适合中医AI研究者与临床辅助系统开发者

传统中医(TCM)具有独特而丰富的知识体系,对大语言模型(LLM)的应用构成挑战。尽管已有针对中医的专用LLM通过监督微调取得进展,但仍面临对齐不足、数据质量差和评估不一致等问题。本研究提出Ladder-base,首个基于组内相对策略优化(GRPO)训练的中医专用大模型。该模型以Qwen2.5-7B-Instruct为基础,仅在TCM-Ladder基准的数据文本子集上训练,其中80%用于训练,剩余20%均分作验证与测试。标准化评估显示,Ladder-base在多项推理指标上优于包括GPT-4、Gemini 2.5、Claude 3、Qwen3在内的通用先进模型,以及BenTsao、HuatuoGPT2、Zhongjing等领域专用模型。结果表明,GRPO是一种有效且高效的策略,可实现大模型在传统医学领域的专家级推理对齐,有助于构建可信、临床可用的中医人工智能系统。

原文摘要 · Abstract (English)

Traditional Chinese Medicine (TCM) presents a rich and structurally unique knowledge system that challenges conventional applications of large language models (LLMs). Although previous TCM-specific LLMs have shown progress through supervised fine-tuning, they often face limitations in alignment, data quality, and evaluation consistency. In this study, we introduce Ladder-base, the first TCM-focused LLM trained with Group Relative Policy Optimization (GRPO), a reinforcement learning method that improves reasoning and factual consistency by optimizing response selection based on intra-group comparisons. Ladder-base is built upon the Qwen2.5-7B-Instruct foundation model and trained exclusively on the textual subset of the TCM-Ladder benchmark, using 80 percent of the data for training and the remaining 20 percent split evenly between validation and test sets. Through standardized evaluation, Ladder-base demonstrates superior performance across multiple reasoning metrics when compared to both state-of-the-art general-purpose LLMs such as GPT-4, Gemini 2.5, Claude 3, and Qwen3 and domain-specific TCM models including BenTsao, HuatuoGPT2, and Zhongjing. These findings suggest that GRPO provides an effective and efficient strategy for aligning LLMs with expert-level reasoning in traditional medical domains and supports the development of trustworthy and clinically grounded TCM artificial intelligence systems.

中医AI强化学习大模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。