arXiv:2501.02497cs.AIcs.CL2025-01综述被引 22

测试时计算让大模型从直觉推理进化为深度思考。

A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

  • 从系统1到系统2,通过反复采样与自我修正提升推理能力。
  • 测试时计算可增强模型对分布偏移的适应性,提升鲁棒性。
  • 适合研究大模型推理机制与智能演进方向的读者。

o1模型在复杂推理任务中的卓越表现表明,测试时计算扩展能进一步释放模型潜力,实现强大的系统2式思维。然而,当前仍缺乏对测试时计算扩展的全面综述。本文追溯测试时计算的概念起源至系统1模型:在系统1中,测试时计算通过参数更新、输入修改、表示编辑和输出校准应对分布偏移,提升鲁棒性与泛化能力;在系统2中,则通过重复采样、自我修正和树搜索增强模型解决复杂问题的推理能力。本综述按系统1向系统2思维演进的脉络组织,强调测试时计算在推动模型从弱系统2迈向强系统2过程中的关键作用,并指出前沿议题与未来方向。

原文摘要 · Abstract (English)

The remarkable performance of the o1 model in complex reasoning demonstrates that test-time compute scaling can further unlock the model's potential, enabling powerful System-2 thinking. However, there is still a lack of comprehensive surveys for test-time compute scaling. We trace the concept of test-time compute back to System-1 models. In System-1 models, test-time compute addresses distribution shifts and improves robustness and generalization through parameter updating, input modification, representation editing, and output calibration. In System-2 models, it enhances the model's reasoning ability to solve complex problems through repeated sampling, self-correction, and tree search. We organize this survey according to the trend of System-1 to System-2 thinking, highlighting the key role of test-time compute in the transition from System-1 models to weak System-2 models, and then to strong System-2 models. We also point out advanced topics and future directions.

推理增强系统2测试时计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。