用六个维度评估大模型推理质量,比只看对错更全面。
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

- 从认知科学出发,设计六维行为评估框架
- 发现局部逻辑与正确性可独立变化,模型表现随上下文波动
- 适合关注模型稳定性和可靠性的人看
尽管大模型在推理基准上取得显著进展,当前评估仍仅依赖最终答案正确性,难以揭示模型推理过程、上下文变化下的行为稳定性及决策效率。本文提出一个统一的多维行为评估框架,基于认知科学定义六项理论维度:正确性(CQ)、一致性(CS)、鲁棒性(RS)、局部逻辑连贯性(LS)、效率(ES)和稳定性(SS)。框架引入部署感知聚合,支持根据上下文选择模型,超越仅凭准确率排名的局限。跨多个大模型与基准的实验揭示了单指标评估掩盖的系统性行为特征,包括局部逻辑连贯性与正确性的正交性、部署上下文相关的排名反转现象,以及小型本地部署模型的非平凡维度表现。判别效度分析表明各维度捕捉的是基本不重叠的信号。该流程为诊断不同部署场景下大模型推理行为提供了基础,领域特定验证是未来方向。
原文摘要 · Abstract (English)
Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes a unified multi-dimensional framework for measuring LLM reasoning quality from a behavioral perspective, operationalizing six theoretically grounded dimensions rooted in cognitive science: Correctness (CQ), Consistency (CS), Robustness (RS), Local Logical Coherence (LS), Efficiency (ES), and Stability (SS). The framework introduces deployment-aware aggregation, enabling context-specific model selection beyond accuracy-based leaderboards. Experiments across multiple LLMs and benchmarks reveal behaviors systematically concealed by single-metric evaluation, including the orthogonality of local logical coherence and correctness, deployment-context-dependent ranking inversions, and non-trivial dimensional profiles in small locally-deployed models. Discriminant validity analysis confirms that the proposed dimensions capture largely non-redundant signals. The resulting pipeline provides a foundation for diagnosing LLM reasoning behavior across deployment contexts, with domain-specific validation as a direction for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。