构建10万条多跳常识推理题,评测大模型长链推理能力。
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
- 基于场景模板生成多跳推理问题,覆盖2-11步推理链。
- 顶尖模型最高仅69.78%准确率,难题集仅47.91%正确。
- 适合评估和优化大模型的常识推理与逻辑一致性能力。
当前大语言模型在长链推理方面仍面临挑战,因自然文本中缺乏充分的显式推理数据。现有基准存在覆盖面窄、推理路径短或构建成本高等问题。我们提出SCoRE(基于场景的常识推理评估),通过实体、关系与逻辑规则的场景模板合成多跳问题,用于评估长链常识推理能力。SCoRE包含10万条中英文双语多项选择题,推理链长度为2至11步,分为不同难度等级。每道题配有细粒度知识标签、明确推理链和难度分级,支持诊断性评估。对o3-mini和Deepseek R1等前沿模型的评估显示,即使最优模型在SCoRE上准确率也仅为69.78%(难题集仅47.91%),错误主要源于罕见知识缺失、逻辑不一致及对简单问题过度解读。SCoRE提供可扩展、可扩展的框架,用于评估、诊断大模型的长链常识推理能力,并指导未来模型设计与训练。
原文摘要 · Abstract (English)
Currently, long-chain reasoning remains a key challenge for large language models (LLMs) because natural texts lack sufficient explicit reasoning data. However, existing benchmarks suffer from limitations such as narrow coverage, short reasoning paths, or high construction costs. We introduce SCoRE (Scenario-based Commonsense Reasoning Evaluation), a benchmark that synthesizes multi-hop questions from scenario schemas of entities, relations, and logical rules to assess long-chain commonsense reasoning. SCoRE contains 100k bilingual (Chinese-English) multiple-choice questions whose reasoning chains span 2-11 hops and are grouped into various difficulty levels. Each question is accompanied by fine-grained knowledge labels, explicit reasoning chains, and difficulty levels for diagnostic evaluation. Evaluation results on cutting-edge LLMs such as o3-mini and Deepseek R1 shows that even the best model attains only 69.78% accuracy on SCoRE (even only 47.91% on the hard set), with errors often stemming from rare knowledge, logical inconsistency, and over-interpretation of simple questions. SCoRE offers a scalable, extensible framework for evaluating and diagnosing the long-chain commonsense reasoning abilities of LLMs and guiding future advances in model design and training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。