提出新方法量化大模型真实推理能力,揭示其依赖记忆而非逻辑推演。
On the Reasoning Capacity of AI Models and How to Quantify It
- 通过位置偏倚实验和两种新模型分析决策机制
- 当前模型在复杂任务中真实推理率低,多靠记忆与模式匹配
- 适合评估大模型可靠性或设计可信应用的开发者
大型语言模型在GPQA和MMLU等基准上表现优异,但在更复杂的推理任务中仍显不足,亟需更严谨的评估方法。本文提出一种现象学框架,超越传统准确率指标,通过多重选择题中的位置偏倚作为案例,利用系统性扰动揭示模型决策的本质。我们构建了概率混合模型(PMM),将模型输出分解为推理、记忆和猜测三部分;同时引入信息论一致性(ITC)分析,量化模型置信度与策略选择间的关系。在控制实验中发现,当前模型的真实推理依然困难,看似成功的表现多源于记忆与模式匹配的复杂组合,而非真正的逻辑演绎。更重要的是,仅靠准确率常夸大模型推理能力,其行为可被表征为认知策略相空间中的动态平衡。该框架提供量化标准,使实际部署能基于策略分布设定可靠性阈值,而非仅依赖整体性能。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have intensified the debate surrounding the fundamental nature of their reasoning capabilities. While achieving high performance on benchmarks such as GPQA and MMLU, these models exhibit limitations in more complex reasoning tasks, highlighting the need for more rigorous evaluation methodologies. We propose a novel phenomenological approach that goes beyond traditional accuracy metrics to probe the underlying mechanisms of model behavior, establishing a framework that could broadly impact how we analyze and understand AI systems. Using positional bias in multiple-choice reasoning tasks as a case study, we demonstrate how systematic perturbations can reveal fundamental aspects of model decision-making. To analyze these behaviors, we develop two complementary phenomenological models: a Probabilistic Mixture Model (PMM) that decomposes model responses into reasoning, memorization, and guessing components and an Information-Theoretic Consistency (ITC) analysis that quantifies the relationship between model confidence and strategy selection. Through controlled experiments on reasoning benchmarks, we show that true reasoning remains challenging for current models, with apparent success often relying on sophisticated combinations of memorization and pattern matching rather than genuine logical deduction. More fundamentally, we demonstrate that accuracy alone often overstates a model's reasoning abilities, as model behavior can be characterized through underlying mechanisms in the phase space of cognitive strategies, revealing how models dynamically balance different approaches when responding to queries. This framework enables quantitative criteria for real-world deployments, allowing applications to specify reliability thresholds based on strategy distributions rather than aggregate performance metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。