提出符号验证框架,提升大模型在物理推理中的测试时扩展能力。
Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset
- 设计符号弱验证框架,利用物理问题结构提升并行推理效率。
- 在TPBench上性能显著优于现有方法,准确率提升超15%。
- 兼顾物理与数学领域,适合需要严谨推理的科学任务研究者。
大型语言模型在复杂推理中展现出强大能力,测试时扩展技术可低成本提升其表现。尽管众多方法已在数学推理基准(如AIME)上得到验证,但这些经验是否适用于高级理论物理领域尚不明确。本文在TPBench物理数据集上评估多种常见测试时扩展方法,并与AIME结果进行对比。为更好利用物理问题的结构特性,我们提出一种新型符号弱验证框架,显著提升并行扩展效果。实证结果表明,该方法在TPBench上显著优于现有方案。同时在AIME上验证了其解决高阶数学问题的有效性。研究揭示了逐步符号验证在应对复杂科学问题中的强大潜力。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown strong capabilities in complex reasoning, and test-time scaling techniques can enhance their performance with comparably low cost. Many of these methods have been developed and evaluated on mathematical reasoning benchmarks such as AIME. This paper investigates whether the lessons learned from these benchmarks generalize to the domain of advanced theoretical physics. We evaluate a range of common test-time scaling methods on the TPBench physics dataset and compare their effectiveness with results on AIME. To better leverage the structure of physics problems, we develop a novel, symbolic weak-verifier framework to improve parallel scaling results. Our empirical results demonstrate that this method significantly outperforms existing test-time scaling approaches on TPBench. We also evaluate our method on AIME, confirming its effectiveness in solving advanced mathematical problems. Our findings highlight the power of step-wise symbolic verification for tackling complex scientific problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。