评测大模型能否提取科学推理中的标准逻辑链条。
ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
- 将复杂推理拆解为皮尔士三种基本推理模式的显式树状结构。
- 10个主流大模型在38,000个观点上表现均存在逻辑完整性和准确性不足。
- 提出新评估指标,揭示模型在科学推理上的系统性缺陷。
大型语言模型在科学领域应用日益广泛,但其通过思维链提示生成的内容通常结构松散、形式非正式,难以判断模型是否真正掌握科学推论的核心逻辑。为此,我们提出新任务Latent Reasoning Chain Extraction (ARCHE),要求模型将复杂推理分解为基于皮尔士三种基本推理模式(演绎、归纳、溯因)的标准逻辑树(RLT)。为支持该任务,我们发布来自70篇《自然·通讯》论文的ARCHE Bench基准,包含超过1,900个引用和38,000个观点。我们设计了两个逻辑感知评估指标:实体覆盖度(EC)衡量内容完整性,推理边准确率(REA)衡量步骤间逻辑有效性。对10个领先大模型的评估显示,模型在REA与EC之间存在权衡,尚无一能完整提取标准推理链。结果表明当前推理模型与科学论证所需的严谨性仍有显著差距。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these outputs are typically unstructured and informal, obscuring whether models truly understand the fundamental reasoning paradigms that underpin scientific inference. To address this, we introduce a novel task named Latent Reasoning Chain Extraction (ARCHE), in which models must decompose complex reasoning arguments into combinations of standard reasoning paradigms in the form of a Reasoning Logic Tree (RLT). In RLT, all reasoning steps are explicitly categorized as one of three variants of Peirce's fundamental inference modes: deduction, induction, or abduction. To facilitate this task, we release ARCHE Bench, a new benchmark derived from 70 Nature Communications articles, including more than 1,900 references and 38,000 viewpoints. We propose two logic-aware evaluation metrics: Entity Coverage (EC) for content completeness and Reasoning Edge Accuracy (REA) for step-by-step logical validity. Evaluations on 10 leading LLMs on ARCHE Bench reveal that models exhibit a trade-off between REA and EC, and none are yet able to extract a complete and standard reasoning chain. These findings highlight a substantial gap between the abilities of current reasoning models and the rigor required for scientific argumentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。