构建多语言常识推理基准,精细评估大模型的思维能力。
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
- 提出细粒度推理技能分类体系,可拆解模型推理过程。
- 在8个主流大模型上测试,高复杂度任务仍显著挑战现有模型。
- 适合研究多语言常识推理与模型可解释性的研究人员使用。
近年来,强化推理能力的大语言模型(LLMs)在复杂推理任务中展现出惊人表现。然而,其如何运用不同人类推理技能的机制仍缺乏深入研究,尤其是在跨语言、跨文化的常识推理方面。为填补这一空白,我们提出了多语言且可扩展的常识推理技能基准(mSCoRe)。该基准包含三个核心组件:(1)新颖的推理技能分类体系,支持对模型推理过程进行细粒度分析;(2)专为常识推理设计的稳健数据生成流程;(3)动态可扩展的难度框架,能随未来模型能力提升而调整任务复杂度。在八个不同规模和训练方式的先进大模型上进行的广泛实验表明,mSCoRe 对当前模型仍具显著挑战性,尤其在高复杂度条件下。结果揭示了现有推理增强型模型在处理多语言通用及文化常识时的局限性。我们进一步对模型推理过程进行了详细分析,指明了未来提升多语言常识推理能力的方向。
原文摘要 · Abstract (English)
Recent advancements in reasoning-reinforced Large Language Models (LLMs) have shown remarkable capabilities in complex reasoning tasks. However, the mechanism underlying their utilization of different human reasoning skills remains poorly investigated, especially for multilingual commonsense reasoning that involves everyday knowledge across different languages and cultures. To address this gap, we propose a \textbf{M}ultilingual and Scalable Benchmark for \textbf{S}kill-based \textbf{Co}mmonsense \textbf{Re}asoning (\textbf{mSCoRe}). Our benchmark incorporates three key components that are designed to systematically evaluate LLM's reasoning capabilities, including: (1) a novel taxonomy of reasoning skills that enables fine-grained analysis of models' reasoning processes, (2) a robust data synthesis pipeline tailored specifically for commonsense reasoning evaluation, and (3) a complexity scaling framework allowing task difficulty to scale dynamically alongside future improvements in LLM abilities. Extensive experiments on eights state-of-the-art LLMs of varying sizes and training approaches demonstrate that \textbf{mSCoRe} remains significantly challenging for current models, particularly at higher complexity levels. Our results reveal the limitations of such reasoning-reinforced models when confronted with nuanced multilingual general and cultural commonsense. We further provide detailed analysis on the models' reasoning processes, suggesting future directions for improving multilingual commonsense reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。