让医疗大模型按需控制推理长度,又准又快。
ControlMed: Adding Reasoning Control to Medical Language Model
- 用细粒度控制标记在推理时调节过程长短。
- 在英韩医疗评测中表现优于或媲美顶尖模型。
- 适合需要平衡准确率与计算效率的临床场景。
增强准确性和可解释性的推理型大语言模型在医疗领域日益受到重视,因其关乎生命安全,临床决策需可靠支持。然而现有模型常生成过长的推理过程,导致显著计算开销和响应延迟,制约其在真实临床环境中的部署。为此,我们提出ControlMed,一种可在推理时通过细粒度控制标记主动调节推理长度的医疗语言模型。该模型采用三阶段训练流程:1)在大规模合成医疗指令数据集上预训练,涵盖直接回答与推理回答;2)使用多长度推理数据及显式长度控制标记进行监督微调;3)通过基于模型的奖励信号强化学习,提升事实准确性与响应质量。在多种英文和韩文医疗基准测试中,模型表现与当前最优水平相当或更优。用户可根据需求灵活权衡推理准确率与计算效率。结果表明,ControlMed是临床问答与医疗信息分析的实用且可适配解决方案。
原文摘要 · Abstract (English)
Reasoning Large Language Models (LLMs) with enhanced accuracy and explainability are increasingly being adopted in the medical domain, as the life-critical nature of clinical decision-making demands reliable support. Despite these advancements, existing reasoning LLMs often generate unnecessarily lengthy reasoning processes, leading to significant computational overhead and response latency. These limitations hinder their practical deployment in real-world clinical environments. To address these challenges, we introduce \textbf{ControlMed}, a medical language model that enables users to actively control the length of the reasoning process at inference time through fine-grained control markers. ControlMed is trained through a three-stage pipeline: 1) pre-training on a large-scale synthetic medical instruction dataset covering both \textit{direct} and \textit{reasoning responses}; 2) supervised fine-tuning with multi-length reasoning data and explicit length-control markers; and 3) reinforcement learning with model-based reward signals to enhance factual accuracy and response quality. Experimental results on a variety of English and Korean medical benchmarks demonstrate that our model achieves similar or better performance compared to state-of-the-art models. Furthermore, users can flexibly balance reasoning accuracy and computational efficiency by controlling the reasoning length as needed. These findings demonstrate that ControlMed is a practical and adaptable solution for clinical question answering and medical information analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。