arXiv:2510.09351cs.CL2025-10ACL被引 2

小模型答对题但推理错,新基准揭示评估盲区

ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering

  • 构建过程级评估基准ReTraceQA,关注推理链合理性
  • 14%-24%题目中小模型答案正确但推理错误
  • 用大模型评判推理过程后,小模型得分下降最高达25%

尽管小型语言模型(SLMs)在常识推理基准上表现日益优异,现有评估方法几乎仅关注最终答案的准确性,忽略了其推理过程的有效性。为此,我们提出ReTraceQA,一个引入过程级评估的新型基准。专家标注的数据集显示,在相当一部分样本中(14%-24%),SLMs虽得出正确答案,但推理过程存在缺陷,表明仅以最终答案对比真值的评估指标常高估其真实能力。事实上,当使用强大型语言模型(LLMs)作为自动化评判者进行推理感知评估时,所有模型和数据集上的性能均显著下降,分数最多降低25%。

原文摘要 · Abstract (English)

While Small Language Models (SLMs) have demonstrated promising performance on an increasingly wide array of commonsense reasoning benchmarks, current evaluation practices rely almost exclusively on the accuracy of their final answers, neglecting the validity of the reasoning processes that lead to those answers. To address this issue, we present ReTraceQA, a novel benchmark that introduces process-level evaluation for commonsense reasoning tasks. Our expert-annotated dataset reveals that in a substantial portion of instances (14-24%), SLMs provide correct final answers despite flawed reasoning processes, suggesting that the capabilities of SLMs are often overestimated by evaluation metrics that focus only on comparing the final answer with the ground truth. Indeed, we show that, when employing strong Large Language Models (LLMs) as automated judges for reasoning-aware evaluation rather than answer-only metrics, SLM performance drops significantly across all models and datasets, with scores decreasing by up to 25%.

常识推理小模型评估推理过程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。