测试时扩展让小模型也能达到顶尖医学推理水平
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
- 通过测试时扩展提升模型推理能力,无需复杂训练
- 4000个推理令牌为最优预算,超量反而降低表现
- 医学知识丰富度比深度推理更重要,适合医疗AI研究者
测试时扩展已成为增强大语言模型推理能力的有效方法。然而,在医学推理领域其效果尚不明确,因医学知识表征与决策过程与数学任务有本质差异。本文首次全面研究测试时扩展在医学推理中的应用,提出m1方法,可在推理阶段显著提升模型医学推理能力。跨多种医学任务的评估显示,测试时扩展持续提升性能,使参数低于100亿的轻量微调模型达到新纪录,320亿模型表现媲美此前700亿级医学大模型。但发现约4000个推理令牌为最佳预算,超出后性能可能下降。预算强制机制虽可促进答案复核,却未必提升整体医学问答准确率,甚至引入错误。案例分析表明,医学知识不足是性能瓶颈。扩大数据规模、提升数据质量、增加模型容量均能强化医学知识基础,推动性能持续提升,尤其在挑战性医学基准上,小模型趋于饱和。研究揭示医学与数学推理的根本差异:仅靠增加推理深度无法发挥测试时扩展优势,必须辅以更丰富的医学知识。
原文摘要 · Abstract (English)
Test-time scaling has emerged as a powerful technique for enhancing the reasoning capabilities of large language models. However, its effectiveness in medical reasoning remains uncertain, as the medical domain fundamentally differs from mathematical tasks in terms of knowledge representation and decision-making processes. In this paper, we provide the first comprehensive investigation of test-time scaling for medical reasoning and present m1, a simple yet effective approach that increases a model's medical reasoning capability at inference. Our evaluation across diverse medical tasks demonstrates that test-time scaling consistently enhances medical reasoning, enabling lightweight fine-tuned models under 10B parameters to establish new state-of-the-art performance, while our 32B model rivals previous 70B-scale medical LLMs. However, we identify an optimal reasoning token budget of approximately 4K, beyond which performance may degrade due to overthinking. Budget forcing, which extends test-time computation through iterative prompts, helps models double-check answers but does not necessarily improve the overall medical QA performance and, in some cases, even introduces errors into previously correct responses. Our case-by-case analysis identifies insufficient medical knowledge as a key bottleneck that prevents further performance gains through test-time scaling. We find that increasing data scale, improving data quality, and expanding model capacity consistently enhance medical knowledge grounding, enabling continued performance improvements, particularly on challenging medical benchmarks where smaller models reach saturation. These findings underscore fundamental differences between medical and mathematical reasoning in LLMs, highlighting that enriched medical knowledge, other than increased reasoning depth alone, is essential for realizing the benefits of test-time scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。