arXiv:2504.18080cs.CLcs.AI2025-04被引 5

优化医学大模型推理稳定性,提升日语医疗问答准确率。

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization

  • 两阶段微调:持续预训练+基于偏好优化的推理增强
  • 日语医学考试准确率达0.868,优于GPT-4o
  • 生成解释时仍保持高精度,避免推理导致性能下降

大型语言模型在医学领域有潜力,但临床应用受限于事实准确性、语言特定局限(如日语)以及生成推理过程的可靠性——这是建立信任的关键。本文提出Preferred-MedLLM-Qwen-72B,一个720亿参数的日语医学专用模型,通过两阶段微调实现高准确率与稳定推理。首先在全面的日语医学语料上进行持续预训练(CPT),注入深度领域知识;其次采用基于偏好的推理偏好优化(RPO),强化可靠推理路径生成,同时保持高答案准确率。在日语医学执照考试基准测试IgakuQA上,该模型达到0.868的准确率,超越GPT-4o(0.866)。关键的是,相较于基线或仅经CPT的模型,在要求生成解释时分别出现最高11.5%和3.8%的准确率下降,本模型仍维持0.868的高水平。这表明RPO有效稳定了推理生成。研究强调在追求准确率的同时,优化可信赖的解释能力至关重要。我们已开源Preferred-MedLLM-Qwen-72B模型权重,推动高风险专业场景下可信大模型的研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show potential in medicine, yet clinical adoption is hindered by concerns over factual accuracy, language-specific limitations (e.g., Japanese), and critically, their reliability when required to generate reasoning explanations -- a prerequisite for trust. This paper introduces Preferred-MedLLM-Qwen-72B, a 72B-parameter model optimized for the Japanese medical domain to achieve both high accuracy and stable reasoning. We employ a two-stage fine-tuning process on the Qwen2.5-72B base model: first, Continued Pretraining (CPT) on a comprehensive Japanese medical corpus instills deep domain knowledge. Second, Reasoning Preference Optimization (RPO), a preference-based method, enhances the generation of reliable reasoning pathways while preserving high answer accuracy. Evaluations on the Japanese Medical Licensing Exam benchmark (IgakuQA) show Preferred-MedLLM-Qwen-72B achieves state-of-the-art performance (0.868 accuracy), surpassing strong proprietary models like GPT-4o (0.866). Crucially, unlike baseline or CPT-only models which exhibit significant accuracy degradation (up to 11.5\% and 3.8\% respectively on IgakuQA) when prompted for explanations, our model maintains its high accuracy (0.868) under such conditions. This highlights RPO's effectiveness in stabilizing reasoning generation. This work underscores the importance of optimizing for reliable explanations alongside accuracy. We release the Preferred-MedLLM-Qwen-72B model weights to foster research into trustworthy LLMs for specialized, high-stakes applications.

医学大模型推理优化日语NLP可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。