用动态交互评估大模型推理过程,揭示其真实思维能力
Evaluating LLM Reasoning Beyond Correctness and CoT

- 引入SIEV框架,通过论点-反论点-综合三阶段评估推理过程
- GPT-5-chat在GSM测试中得分下降超40分,暴露推理脆弱性
- 适合关注模型思维深度与可信度的研究者和开发者
当前大模型评估仅关注答案正确性,却忽略了推理过程。本文提出SIEV框架,基于辩证哲学思想,通过显式的论点-反论点-综合互动来评估推理的动态轨迹。该框架能揭示模型对质疑的抗性、冲突下的适应性以及多视角整合能力,这些是传统正确率指标无法捕捉的维度。在GSM和MMLU数据集上的实验证明,顶尖模型如GPT-5-chat在采用过程导向评估时,于GSM上得分下降超过40分(满分100)。这一转变使我们能区分结构化推理与表面模式生成,为理解与评估大模型的真实推理能力提供了更透明、更严谨的基础。
原文摘要 · Abstract (English)
What does it truly mean for a language model to "reason"? Current evaluations reward models' correct standalone answers-but correctness alone reveals little about the process that produced them. We argue that reasoning should be understood not as a static chain of steps but as a dynamic trajectory in which ideas interact, clash, and evolve into integrated insights. Building on the philosophical tradition of dialectics, we introduce SIEV, a structured evaluation framework that assesses reasoning through explicit thesis-antithesis-synthesis interactions. SIEV produces interpretable trajectories that highlight key properties of reasoning-robustness to challenge, adaptability under conflict, and synthesis across competing viewpoints-dimensions that conventional correctness-based metrics cannot capture. Empirical results on GSM and MMLU demonstrate substantial gaps in the reasoning abilities of state-of-the-art models: for example, GPT-5-chat loses more than 40 points (out of 100) on GSM when evaluated through SIEV's process-oriented lens. By shifting focus from what answer a model gives to how it arrives there, SIEV enables a more transparent and principled distinction between structured reasoning and surface-level pattern generation offering a clearer foundation for assessing and understanding the reasoning capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。