好推理链能用更少参数完成任务,提升模型泛化能力。
Effective Reasoning Chains Reduce Intrinsic Dimensionality
- 用内在维度衡量推理链效率,越低表示越高效
- 在GSM8K上,有效推理链使内在维度降低30%以上
- 适合研究模型泛化与推理机制的学者参考
链式思维(CoT)及其变体显著提升了语言模型在复杂推理任务上的表现,但不同策略如何促进泛化仍不明确。本文提出以内在维度作为量化指标,衡量完成任务所需的最小模型维度。在固定模型架构下,通过不同推理策略改变任务形式,发现有效策略可显著降低任务内在维度。在GSM8K数据集上使用Gemma-3 1B和4B模型验证,内在维度越低,模型在分布内和分布外数据上的泛化性能越好,二者呈强负相关。结果表明,高效推理链通过更优压缩任务信息,用更少参数实现学习,为分析推理过程提供了新量化工具。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, yet the precise mechanisms by which different strategies facilitate generalization remain poorly understood. While current explanations often point to increased test-time computation or structural guidance, establishing a consistent, quantifiable link between these factors and generalization remains challenging. In this work, we identify intrinsic dimensionality as a quantitative measure for characterizing the effectiveness of reasoning chains. Intrinsic dimensionality quantifies the minimum number of model dimensions needed to reach a given accuracy threshold on a given task. By keeping the model architecture fixed and varying the task formulation through different reasoning strategies, we demonstrate that effective reasoning strategies consistently reduce the intrinsic dimensionality of the task. Validating this on GSM8K with Gemma-3 1B and 4B, we observe a strong inverse correlation between the intrinsic dimensionality of a reasoning strategy and its generalization performance on both in-distribution and out-of-distribution data. Our findings suggest that effective reasoning chains facilitate learning by better compressing the task using fewer parameters, offering a new quantitative metric for analyzing reasoning processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。