arXiv:2507.01334cs.CL2025-07被引 2

探究大模型解物理题时用符号还是数值,发现符号推导更有效。

Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs

  • 用符号推导而非直接计算,提升推理准确性。
  • 在SciBench基准上达到顶尖水平,正确率显著高于以往方法。
  • 少量示例提示可进一步提升性能,适合研究物理推理的学者。

解决物理推理问题长期是大语言模型(LLMs)的挑战,需要深刻的概念理解与精巧的问题求解技巧。本研究考察了Deepseek-R1等先进指令微调推理模型在SciBench基准中多样物理问题上的表现。全面实验表明,这些推理模型不仅在复杂物理问题上达到最先进的准确率,还展现出以符号推导为核心的独特推理模式。此外,即使对于高度先进的模型,引入少量示例提示仍能带来可测量的准确率提升,凸显持续优化潜力。

原文摘要 · Abstract (English)

Navigating the complexities of physics reasoning has long been a difficult task for Large Language Models (LLMs), requiring a synthesis of profound conceptual understanding and adept problem-solving techniques. In this study, we investigate the application of advanced instruction-tuned reasoning models, such as Deepseek-R1, to address a diverse spectrum of physics problems curated from the challenging SciBench benchmark. Our comprehensive experimental evaluation reveals the remarkable capabilities of reasoning models. Not only do they achieve state-of-the-art accuracy in answering intricate physics questions, but they also generate distinctive reasoning patterns that emphasize on symbolic derivation. Furthermore, our findings indicate that even for these highly sophisticated reasoning models, the strategic incorporation of few-shot prompting can still yield measurable improvements in overall accuracy, highlighting the potential for continued performance gains.

物理推理符号推导大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。