通过延长推理时间,医学大模型在诊断与治疗决策上表现显著提升。
O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
- 在500样本小训练集上,通过延长推理时间提升性能6%-11%。
- 任务越复杂,所需推理链条越长,证明深度思考对难题必要。
- 生成的鉴别诊断符合假说演绎法,系统缩小病因范围。
继前期关于O1复现的研究(第一部分:旅程学习[秦等,2024],第二部分:蒸馏[黄等,2024])之后,本文探索了大型语言模型在医学推理任务中推理时长扩展的潜力,涵盖从诊断决策到治疗规划的多种场景。在不同复杂度的医学基准测试(MedQA、Medbullets、JAMA临床挑战)上进行大量实验后发现:(1)增加推理时间确实带来性能提升,在仅500个样本的小型训练集下,性能提升达6%-11%;(2)任务复杂性与推理链长度呈正相关,证实复杂问题需要更长的思维过程;(3)模型生成的鉴别诊断遵循假说演绎法,列出可能病因并逐步依据证据排除。这些结果表明,推理时长扩展与旅程学习相结合,有望显著增强大模型在真实临床场景中的推理能力。
原文摘要 · Abstract (English)
Building upon our previous investigations of O1 replication (Part 1: Journey Learning [Qin et al., 2024] and Part 2: Distillation [Huang et al., 2024]), this work explores the potential of inference-time scaling in large language models (LLMs) for medical reasoning tasks, ranging from diagnostic decision-making to treatment planning. Through extensive experiments on medical benchmarks of varying complexity (MedQA, Medbullets, and JAMA Clinical Challenges), our investigation reveals several key insights: (1) Increasing inference time does lead to improved performance. With a modest training set of 500 samples, our model yields substantial performance improvements of 6%-11%. (2) Task complexity directly correlates with the required length of reasoning chains, confirming the necessity of extended thought processes for challenging problems. (3) The differential diagnoses generated by our model adhere to the principles of the hypothetico-deductive method, producing a list of potential conditions that may explain a patient's symptoms and systematically narrowing these possibilities by evaluating the evidence. These findings demonstrate the promising synergy between inference-time scaling and journey learning in advancing LLMs' real-world clinical reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。