arXiv:2602.02497cs.CLcs.AI2026-02被引 1

构建双轴诊断框架,精准分析大模型在理工科推理中的短板。

STEMVerse: A Dual-Axis Diagnostic Framework for STEM Reasoning in Large Language Models

  • 从学科与认知深度双维度重新标注超2万道理工题
  • 发现大模型在特定学科和复杂推理上存在结构性缺陷
  • 适合研究者定位模型弱点,指导模型优化方向

随着大语言模型(LLMs)在复杂推理任务中取得显著进展,评估其在科学、技术、工程和数学(STEM)领域的表现成为衡量机器智能的核心方法。然而,现有评估范式常将基准视为孤立的‘信息孤岛’,仅提供单一的整体得分,忽视了学科专精与认知深度的复杂性。这种结果导向的方法无法区分模型错误是源于领域知识不足,还是认知能力欠缺,限制了诊断价值。为此,我们提出STEMVerse,一个系统化分析LLMs STEM推理能力的诊断框架。该框架通过学术专精与认知复杂度双重维度,映射推理所需能力。我们将主流基准中的20,000余道STEM问题重新整合至统一的‘学科×认知’能力空间,并为每道题分配双轴标签。基于此框架,我们系统评估了不同参数规模与训练范式的代表性LLM家族。实证结果揭示了STEM推理中的结构性失败模式。通过融合多学科覆盖与细粒度认知分层,STEMVerse为理解大模型的科学推理特征提供了清晰且可操作的视角。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) achieve significant breakthroughs in complex reasoning tasks, evaluating their proficiency in science, technology, engineering, and mathematics (STEM) has become a primary method for measuring machine intelligence. However, current evaluation paradigms often treat benchmarks as isolated "silos," offering only monolithic aggregate scores that neglect the intricacies of both academic specialization and cognitive depth. This result-oriented approach fails to distinguish whether model errors stem from insufficient domain knowledge or deficiencies in cognitive capacity, thereby limiting the diagnostic value. To address this, we propose STEMVerse, a diagnostic framework designed to systematically analyze the STEM reasoning capabilities of LLMs. This framework characterizes model performance across academic specialization and cognitive complexity to map the capability required for reasoning. We re-aggregate over 20,000 STEM problems from mainstream benchmarks into a unified "Discipline $\times$ Cognition" capability space, assigning dual-axis labels to every instance. Utilizing this unified diagnostic framework, we systematically evaluate representative LLM families across varying parameter scales and training paradigms. Our empirical results reveal structural failure patterns in STEM reasoning. By integrating multi-disciplinary coverage and fine-grained cognitive stratification into a unified framework, STEMVerse provides a clear and actionable perspective for understanding the scientific reasoning characteristics of LLMs.

大模型评估认知诊断理工推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。