arXiv:2601.22446cs.AI2026-01中稿 · ICML被引 2

提出B-PAC推理方法,实现高效安全的在线决策,误差可控。

Anytime Safe PAC Efficient Reasoning

  • 用反倾向评分构建统计检验,动态调整推理路由阈值
  • 实验显示思考模型使用率降低81.01%,性能损失可控在设定范围内
  • 适合对安全性和效率都有要求的实时推理场景

大型推理模型(LRMs)在复杂任务上表现优异,但计算成本高、延迟大。现有选择性思考策略虽能通过将简单问题转给非思考模型提升效率,但在在线环境下常导致不可控错误,因非思考模型的性能损失仅部分可观测且数据分布不稳。为此,我们提出贝叶斯概率近似正确(B-PAC)推理,一种支持任意时间安全与高效在线推理的方法。具体地,利用反倾向评分估计器构建候选阈值的检验超鞅,并根据累积统计证据动态调整路由阈值。理论上,我们建立了任意时间有效的性能损失控制和推理效率保证。大量实验表明,B-PAC推理显著降低计算开销,思考模型使用率最高下降81.01%,同时将性能损失控制在用户指定水平以下。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex tasks but suffer from high computational costs and latency. While selective thinking strategies improve efficiency by routing easy queries to non-thinking models, existing approaches often incur uncontrollable errors, especially in online settings where the performance loss of a non-thinking model is only partially observed and data are non-stationary. To address this, we propose Betting Probably Approximately Correct (B-PAC) reasoning, a principled method that enables anytime safe and efficient online reasoning under partial feedback. Specifically, we utilize inverse propensity scoring estimators to construct test supermartingales for candidate thresholds, and then dynamically adjust the routing threshold based on the accumulated statistical evidence of safety. Theoretically, we establish the anytime-valid performance loss control and the efficiency of B-PAC reasoning. Extensive experiments demonstrate that B-PAC reasoning significantly reduces computational overhead, decreasing thinking model usage by up to 81.01\%, while controlling the performance loss below the user-specified level.

推理优化在线学习安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。