arXiv:2506.04611cs.CL2025-06综述被引 6

提升大模型推理效率,通过增强生成多样性实现更优测试时扩展效果

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning

  • 提出基于多样性感知的前缀微调方法ADAPT,优化推理过程
  • 在数学推理任务中仅用1/8计算量达到80%准确率
  • 适合追求高效推理且关注输出多样性的研究者

测试时扩展(TTS)通过在推理阶段增加计算资源,提升大语言模型的推理性能。本文系统梳理了TTS方法,将其分为采样类、搜索类和轨迹优化类策略。研究发现,为推理优化的模型往往生成结果多样性不足,限制了TTS效果。为此,提出ADAPT(一种多样性感知的前缀微调方法),采用以多样性为导向的数据策略进行轻量级前缀微调。在数学推理任务上的实验表明,ADAPT在仅使用强基线1/8计算量的情况下,达到80%的准确率。研究强调生成多样性对最大化TTS效果的关键作用。

原文摘要 · Abstract (English)

Test-Time Scaling (TTS) improves the reasoning performance of Large Language Models (LLMs) by allocating additional compute during inference. We conduct a structured survey of TTS methods and categorize them into sampling-based, search-based, and trajectory optimization strategies. We observe that reasoning-optimized models often produce less diverse outputs, which limits TTS effectiveness. To address this, we propose ADAPT (A Diversity Aware Prefix fine-Tuning), a lightweight method that applies prefix tuning with a diversity-focused data strategy. Experiments on mathematical reasoning tasks show that ADAPT reaches 80% accuracy using eight times less compute than strong baselines. Our findings highlight the essential role of generative diversity in maximizing TTS effectiveness.

测试时扩展推理优化多样性增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。