arXiv:2506.00189cs.AIcs.CL2025-06被引 1

让大模型推理过程可调控,避免想得少或想得多

Control-R: Towards controllable test-time scaling

  • 通过结构化控制信号引导树搜索式推理
  • 在AIME2024和MATH500上达32B规模最先进水平
  • 适合需要灵活调整推理强度的复杂任务场景

本文针对大推理模型在长链式思维(L-CoT)中存在思考不足与过度思考的问题,提出测试时的推理控制场(RCF)——一种基于树搜索视角的新型控制方法。该方法通过注入结构化控制信号,使模型在解决复杂任务时可根据给定条件动态调整推理强度。我们构建了Control-R-4K数据集,包含需详细推理过程及对应控制字段的挑战性问题。为进一步提升控制能力,提出条件蒸馏微调(CDF)方法,训练模型(特别是Control-R-32B)在测试阶段有效调节推理力度。在AIME2024和MATH500等基准上的实验表明,该方法在32B规模下达到当前最优性能,并实现可控的长链式推理过程。整体工作提出了一种高效的可调控测试时扩展推理范式。

原文摘要 · Abstract (English)

This paper target in addressing the challenges of underthinking and overthinking in long chain-of-thought (CoT) reasoning for Large Reasoning Models (LRMs) by introducing Reasoning Control Fields (RCF)--a novel test-time approach that injects structured control signals to guide reasoning from a tree search perspective. RCF enables models to adjust reasoning effort according to given control conditions when solving complex tasks. Additionally, we present the Control-R-4K dataset, which consists of challenging problems annotated with detailed reasoning processes and corresponding control fields. To further enhance reasoning control, we propose a Conditional Distillation Finetuning (CDF) method, which trains model--particularly Control-R-32B--to effectively adjust reasoning effort during test time. Experimental results on benchmarks such as AIME2024 and MATH500 demonstrate that our approach achieves state-of-the-art performance at the 32B scale while enabling a controllable Long CoT reasoning process (L-CoT). Overall, this work introduces an effective paradigm for controllable test-time scaling reasoning.

推理控制链式思维测试时扩展大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。