arXiv:2512.24063cs.LG2025-12被引 7

拆解大模型推理能力,揭示强化学习为何比微调更抗退化。

How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns

  • 按计算、检索等原子技能拆解推理,构建细粒度评估框架。
  • 强化学习模型在数学与科学任务中保持更稳定的推理能力,微调模型易退化。
  • 适合关注模型泛化机制、训练策略设计的研究者阅读。

大型语言模型(LLMs)表现出显著不同的泛化行为:监督微调(SFT)常导致能力收缩,而强化学习(RL)微调则倾向于保留能力。现有研究多依赖粗粒度准确率指标,难以解释这一差异。本文提出一种新基准,将推理分解为计算、事实检索、模拟、枚举和诊断等原子核心技能,构建了分析大模型推理本质的可操作框架。通过隔离并测量这些核心技能,该基准揭示了特定认知能力在后训练阶段的涌现、迁移与崩溃规律。结合对低层统计模式(如分布偏移、参数统计)的分析,实现了对数学、科学推理及非推理任务中泛化演化的细粒度研究。元探针框架追踪不同训练阶段的模型行为,发现RL微调模型维持更稳定的行为特征,推理能力更抗退化;而SFT模型表现出更剧烈的漂移并过拟合表层模式。本工作深化了对大模型推理本质的理解,并为设计促进广泛、鲁棒泛化的训练策略提供了新思路。

原文摘要 · Abstract (English)

Large Language Models (LLMs) display strikingly different generalization behaviors: supervised fine-tuning (SFT) often narrows capability, whereas reinforcement-learning (RL) tuning tends to preserve it. The reasons behind this divergence remain unclear, as prior studies have largely relied on coarse accuracy metrics. We address this gap by introducing a novel benchmark that decomposes reasoning into atomic core skills such as calculation, fact retrieval, simulation, enumeration, and diagnostic, providing a concrete framework for addressing the fundamental question of what constitutes reasoning in LLMs. By isolating and measuring these core skills, the benchmark offers a more granular view of how specific cognitive abilities emerge, transfer, and sometimes collapse during post-training. Combined with analyses of low-level statistical patterns such as distributional divergence and parameter statistics, it enables a fine-grained study of how generalization evolves under SFT and RL across mathematical, scientific reasoning, and non-reasoning tasks. Our meta-probing framework tracks model behavior at different training stages and reveals that RL-tuned models maintain more stable behavioral profiles and resist collapse in reasoning skills, whereas SFT models exhibit sharper drift and overfit to surface patterns. This work provides new insights into the nature of reasoning in LLMs and points toward principles for designing training strategies that foster broad, robust generalization.

大模型推理泛化能力强化学习训练机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。