arXiv:2602.17544cs.AIcs.CL2026-02被引 3

评估推理过程质量,提出可复用性和可验证性新指标

Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability

  • 分离思考与执行,用框架测试推理的可复用性
  • 专用推理模型的推理质量不比通用大模型高
  • 现有准确率榜单无法反映真实推理能力

在搜索与排序等多智能体信息检索流程中,基于大语言模型的智能体通过思维链(Chain-of-Thought, CoT)交换中间推理。当前评估仅关注任务最终准确率,忽略了推理过程本身的质量。为此,本文提出两个新指标:可复用性与可验证性。采用Thinker-Executor框架将CoT生成与执行解耦。可复用性衡量执行者重用思考者推理的能力,可验证性衡量执行者依据推理内容匹配原答案的频率。我们在五个基准上,用四个思考者模型与十名执行者模型进行评估。结果表明,可复用性与可验证性与标准准确率无相关性,揭示了当前以准确率为核心的排行榜在评估推理能力上的盲区。令人意外的是,专用推理模型的思维链并不总是比Llama、Gemma等通用大模型更具可复用性或可验证性。

原文摘要 · Abstract (English)

In multi-agent IR pipelines for tasks such as search and ranking, LLM-based agents exchange intermediate reasoning in terms of Chain-of-Thought (CoT) with each other. Current CoT evaluation narrowly focuses on target task accuracy. However, this metric fails to assess the quality or utility of the reasoning process itself. To address this limitation, we introduce two novel measures: reusability and verifiability. We decouple CoT generation from execution using a Thinker-Executor framework. Reusability measures how easily an Executor can reuse the Thinker's CoT. Verifiability measures how frequently an Executor can match the Thinker's answer using the CoT. We evaluated four Thinker models against a committee of ten Executor models across five benchmarks. Our results reveal that reusability and verifiability do not correlate with standard accuracy, exposing a blind spot in current accuracy-based leaderboards for reasoning capability. Surprisingly, we find that CoTs from specialized reasoning models are not consistently more reusable or verifiable than those from general-purpose LLMs like Llama and Gemma.

思维链评估指标多智能体推理质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。