提升大模型物理推理逻辑性,让答案更合理可信。
Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics

- 构建评估逻辑性的标准与引导训练的数据采样方法。
- 在三种大模型上验证,逻辑性提升显著改善解题能力。
- 适合关注科学推理可靠性与可解释性的研究者。
随着大语言模型(LLMs)推理能力的持续提升,其在科学推理任务中的应用受到广泛关注。现有研究主要通过在更大、更全面的数据集上训练并延长推理链来提升模型在科学问答基准上的表现,但忽略了科学推理的本质——逻辑性,即确保推理步骤有效以得出可靠结论的理性基础。本文首次系统探究了大模型科学推理中的内在逻辑性,并提出一种科学逻辑性增强方法,包括一套评估标准和基于逻辑性引导的数据采样方法,以提升推理的逻辑忠实度与任务性能。进一步以具有多样化逻辑结构与形式化的物理学为例,从学术文献中提取科学问题,构建高质量高逻辑性数据集。基于三种不同骨干模型的实验表明:1)所构建的训练数据能有效提升大模型推理的科学逻辑性;2)增强的科学逻辑性在解决科学问题中起关键作用。代码已开源。
原文摘要 · Abstract (English)
With the continuous advancement of reasoning abilities in Large Language Models (LLMs), their application to scientific reasoning tasks has gained significant research attention. Current research primarily emphasizes boosting LLMs' performance on scientific QA benchmarks by training on larger, more comprehensive datasets with extended reasoning chains. However, these approaches neglect the essence of the scientific reasoning process -- logicality, which is the rational foundation to ensure the validity of reasoning steps leading to reliable conclusions. In this work, we make the first systematic investigation into the internal logicality underlying LLM scientific reasoning, and develop a scientific logicality-enriched methodology, including a set of assessment criteria and data sampling methods for logicality-guided training, to improve the logical faithfulness as well as task performance. Further, we take physics, characterized by its diverse logical structures and formalisms, as an exemplar discipline to practise the above methodology. For data construction, we extract scientific problems from academic literature and sample a high-quality dataset exhibiting strong logicality. Experiments based on three different backbone LLMs reveal that: 1) the training data we constructed can effectively improve the scientific logicality in LLM reasoning; and 2) the enriched scientific logicality plays a critical role in solving scientific problems. Code is available at \href{https://github.com/ScienceOne-AI/PhysLogic}{https://github.com/ScienceOne-AI/PhysLogic}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。