发现大模型推理过程可压缩近半,核心步骤仅需3步就占65%支持力。
Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces

- 从冗长推理链中提取最小必要步骤集,量化冗余与必要性。
- 平均46%步骤可移除,86%情况下答案不变,前3步贡献65%推理支撑。
- 压缩后推理更清晰,跨模型迁移准确率高达85%。
语言模型生成的思维链常包含多余步骤,我们通过过完备推理链研究其必要性。定义最小核心为维持最终答案或预测分布的最小步骤子集,提出压缩比、冗余质量、步骤必要性和必要性集中度等指标。在六项涵盖算术、竞赛数学、专家科学推理和常识多跳问答的基准上,发现显著过完备:平均46%的步骤可在贪婪提取下移除,86%案例保持原答案。同时,预测支持高度集中,前3步平均贡献65%必要性质量。最小核心使推理几何更清晰:相比完整链,正确与错误链分离度提升11点,内在维度降低34%,跨模型族迁移时非对角线答案保留率达85%。理论上证明了最小充分子集存在性、贪心消除的局部不可约性,以及过完备与稀疏必要性的验证证书。结果表明,完整推理链常冗长过载,而最小核心揭示了语言模型预测的真实支撑结构。
原文摘要 · Abstract (English)
Language models often generate long chain-of-thought traces, but it remains unclear how much of this reasoning is necessary for preserving the final prediction. We study this through the lens of overcomplete reasoning traces: generated traces that contain more intermediate steps than are needed to support the model's answer. We define the minimal core as the smallest subset of steps that preserves either the final answer or predictive distribution, and introduce metrics for compression ratio, redundancy mass, step necessity, and necessity concentration. Across six deliberative reasoning benchmarks spanning arithmetic, competition mathematics, expert scientific reasoning, and commonsense multi-hop QA, we find substantial overcompleteness: on average, 46% of steps are removable under greedy minimal-core extraction while preserving the original answer in 86% of cases. We also find that predictive support is concentrated: the top three steps account for 65% of measured necessity mass on average. Beyond compression, minimal cores expose a cleaner geometry of reasoning: compared with full traces, they improve correct-incorrect trace separation by 11 points, reduce estimated intrinsic dimensionality by 34%, and transfer across model families with 85% off-diagonal answer retention. Theoretically, we establish existence of minimal sufficient subsets, local irreducibility guarantees for greedy elimination, and certificates of overcompleteness and sparse necessity. Together, these results suggest that full reasoning traces are often verbose and overcomplete, while minimal cores isolate the effective support underlying language-model predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。