arXiv:2602.09555cs.CL2026-02被引 2

动态分配推理计算量,让语言模型更聪明地思考。

Adaptive Test-Time Compute Allocation for Block Diffusion Language Models in Complex Reasoning

  • 根据模型信心动态调整去噪强度,加速推理
  • 大块粗思、小块精修,提升复杂推理效率2.38倍
  • 适合需要高效长链推理的应用场景

块扩散语言模型在推理任务中表现出色且可扩展,但其测试阶段的计算分配尚未深入研究,导致长链思维推理中存在速度与效果的权衡难题。为此,我们提出统一的测试时计算分配框架,实现逐步解码与块级生成的自适应。在解码层面,提出有界自适应置信度解码(BACD),一种基于难度感知的采样策略,动态调整去噪过程以加速推理并控制误差累积。超越逐步自适应,引入‘粗思细评’(TCCF)范式,用大块进行探索性思考,小块进行精确修正。为应对不同块大小配置下的训练不稳定性,采用渐进式块大小扩展策略,缓解大块时性能下降问题。在六个基准上的大量实验表明,采用BACD与TCCF的TDAR-8B模型相比强基线TraDo-8B,在平均准确率提升3.4%的同时,推理速度提升2.38倍,充分释放了块扩散模型在复杂推理中的潜力。

原文摘要 · Abstract (English)

Recent advances in block diffusion language models have demonstrated competitive performance and strong scalability on reasoning tasks. However, their test-time compute allocation remains largely unexplored, leaving a critical speed-effectiveness trade-off unresolved in long Chain-of-Thought reasoning. To address this, we propose a unified test-time compute allocation framework that introduces adaptivity in both step-wise decoding and blockwise generation. At the decoding level, we propose Bounded Adaptive Confidence Decoding (BACD), a difficulty-aware sampling strategy that dynamically adjusts denoising based on model confidence, accelerating inference while controlling error accumulation. Beyond step-wise adaptivity, we introduce the Think Coarse, Critic Fine (TCCF) paradigm that allocates large block sizes for exploratory thinking and smaller block sizes for precise refinement. To stabilize training under varying block configurations, we adopt Progressive Block Size Extension, which mitigates quality degradation when scaling up block sizes. Extensive evaluations on six benchmarks show that our TDAR-8B model with BACD and TCCF achieves a 2.38$\times$ speedup and +3.4% average accuracy over the strong TraDo-8B baseline, unlocking the potential of block diffusion in complex reasoning.

扩散模型推理优化自适应计算复杂推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。