发现推理模型内部的熵梯度反转现象,提升逻辑推理能力。
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models

- 通过熵与梯度负相关性揭示模型推理机制
- 新方法在多任务上超越现有最优模型
- 适合研究大模型推理原理与强化学习优化者
大型推理模型(LRMs)推动了从快速文本生成到系统化逐步推理的范式转变,在复杂数学与逻辑任务中达到顶尖性能。然而,当前领域存在两大挑战:一是词元层面行为分析与内部推理机制之间的根本鸿沟;二是依赖昂贵外部验证器的强化学习(RL)在推理优化中存在不稳定性。本文首次识别并形式化定义了「熵-梯度反演」——一种词元熵与逻辑梯度间的稳健负相关关系,可作为大型推理模型推理能力的几何指纹。基于此,提出「相关性正则化组策略优化」(CorR-PO),将该反演特征嵌入强化学习奖励正则化中。在多个模型规模和多种推理基准上的大量实验表明,CorR-PO持续优于现有最先进基线,证实更强的反演程度直接对应更优的推理表现。
原文摘要 · Abstract (English)
The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive ``fast thinking'' text generation to systematic, step-by-step ``slow thinking'' reasoning, unlocking state-of-the-art performance in complex mathematical and logical tasks. However, the field faces \textit{the fundamental gap between token-level behavioral analysis and internal reasoning mechanisms, and the instability of reinforcement learning (RL) for reasoning optimization relying on costly external verifiers}. We identify and formally define \textbf{Entropy-Gradient Inversion}, a robust negative correlation between token entropy and logit gradients that acts as a definitive geometric fingerprint for LRM reasoning capability. Building on this, we propose \textbf{Correlation-Regularized Group Policy Optimization (CorR-PO)}, which embeds this inversion signature into RL reward regularization. Extensive experiments on various reasoning benchmarks across multiple model scales show CorR-PO consistently outperforms state-of-the-art baselines, confirming that stronger inversion directly correlates with superior reasoning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。