arXiv:2508.12140cs.CL2025-08被引 5

首次揭示医疗推理中计算资源与推理质量的对数关系。

Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality

  • 通过控制思考预算,发现准确率随预算和模型大小呈对数增长。
  • 小模型在扩展思考时提升更明显,15-20%优于大模型的5-10%。
  • 提出三类效率模式,适配实时、常规与关键诊断场景。

本研究首次系统评估了医疗推理任务中思考预算机制,揭示了计算资源与推理质量间的根本缩放规律。在涵盖15个不同专科与难度等级的医学数据集上,对Qwen3(1.7B至235B参数)和DeepSeek-R1(1.5B至70B参数)两大模型家族进行了测试。通过控制思考预算从0到无限令牌,建立对数缩放关系:准确率提升遵循可预测模式。识别出三种效率范式:高效率(0–256令牌),适用于实时应用;平衡型(256–512令牌),提供最佳性价比;高精度型(>512令牌),仅适用于关键诊断。值得注意的是,小模型在扩展思考时获益显著,相对提升达15–20%,远超大模型的5–10%,表明思考预算对容量受限模型更具互补价值。神经科与消化科需更深层推理,而心血管与呼吸科需求较低。Qwen3原生接口与针对DeepSeek-R1提出的截断法结果一致,验证了思考预算概念的跨架构通用性。这些发现确立思考预算控制是优化医疗AI系统的核心机制,支持动态资源分配并保障医疗部署所需的透明性。

原文摘要 · Abstract (English)

This study presents the first comprehensive evaluation of thinking budget mechanisms in medical reasoning tasks, revealing fundamental scaling laws between computational resources and reasoning quality. We systematically evaluated two major model families, Qwen3 (1.7B to 235B parameters) and DeepSeek-R1 (1.5B to 70B parameters), across 15 medical datasets spanning diverse specialties and difficulty levels. Through controlled experiments with thinking budgets ranging from zero to unlimited tokens, we establish logarithmic scaling relationships where accuracy improvements follow a predictable pattern with both thinking budget and model size. Our findings identify three distinct efficiency regimes: high-efficiency (0 to 256 tokens) suitable for real-time applications, balanced (256 to 512 tokens) offering optimal cost-performance tradeoffs for routine clinical support, and high-accuracy (above 512 tokens) justified only for critical diagnostic tasks. Notably, smaller models demonstrate disproportionately larger benefits from extended thinking, with 15 to 20% improvements compared to 5 to 10% for larger models, suggesting a complementary relationship where thinking budget provides greater relative benefits for capacity-constrained models. Domain-specific patterns emerge clearly, with neurology and gastroenterology requiring significantly deeper reasoning processes than cardiovascular or respiratory medicine. The consistency between Qwen3 native thinking budget API and our proposed truncation method for DeepSeek-R1 validates the generalizability of thinking budget concepts across architectures. These results establish thinking budget control as a critical mechanism for optimizing medical AI systems, enabling dynamic resource allocation aligned with clinical needs while maintaining the transparency essential for healthcare deployment.

医疗推理思考预算缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。