arXiv:2510.09133cs.AIcs.LG2025-10被引 3

提出可保证性能损失的高效推理方法,动态切换思考与非思考模式。

On the Provable Performance Guarantee of Efficient Reasoning Models

  • 基于置信上界动态决定是否切换至低耗推理模式。
  • 在多个基准测试中节省计算资源,且性能损失可控。
  • 适合对可靠性要求高的高风险应用场景。

大型推理模型(LRMs)在复杂问题求解任务中取得显著进展,但部署时计算成本高。为提升效率,一种实用方向是动态切换模型的思考与非思考模式。然而,此类方法常引入额外推理误差,且缺乏性能损失的统计保障,这对高风险应用至关重要。本文提出概率近似正确(PAC)推理,可在用户指定容差下控制性能损失。具体地,构建性能损失的置信上界,并确定切换阈值。理论上,使用该阈值在思考与非思考模式间切换,可实现分布无关的性能损失有界。在多个推理基准上的实验表明,该方法能有效节省计算开销,并严格控制用户指定的性能损失。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have achieved remarkable progress in complex problem-solving tasks. Despite this success, LRMs typically suffer from high computational costs during deployment, highlighting a need for efficient inference. A practical direction of efficiency improvement is to switch the LRM between thinking and non-thinking modes dynamically. However, such approaches often introduce additional reasoning errors and lack statistical guarantees for the performance loss, which are critical for high-stakes applications. In this work, we propose Probably Approximately Correct (PAC) reasoning that controls the performance loss under the user-specified tolerance. Specifically, we construct an upper confidence bound on the performance loss and determine a threshold for switching to the non-thinking model. Theoretically, using the threshold to switch between the thinking and non-thinking modes ensures bounded performance loss in a distribution-free manner. Our comprehensive experiments on reasoning benchmarks show that the proposed method can save computational budgets and control the user-specified performance loss.

推理模型效率优化性能保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。