arXiv:2608.01835cs.AI2026-08

通过几何分析揭示微调与强化学习如何不同地改变语言模型行为。

Rewriting or Reweighting? A Geometric Account in Language Models

论文配图:Rewriting or Reweighting? A Geometric Account in Language Models
图 1 · 摘自论文原文
  • 用行为流形分析提取模型中的关键行为坐标,构建低维局部图表。
  • 监督微调显著重写行为几何,而强化学习主要调整权重不改变基础结构。
  • 适合研究模型机制、对齐偏差或可控生成的开发者和研究人员。

后训练可显著改变语言模型的行为,但总体行为指标无法区分是移除旧机制、创建新机制,还是改变已有机制的使用方式。本文通过两种机制迥异的失败现象——解码时重复成瘾和迎合偏好——展开研究。提出行为流形分析方法,通过选择稀疏的行为关联坐标,并将其映射到低维局部图表中,分离出行为特异性几何结构。在两个互补空间中构建图表:ACT捕捉运行时激活状态,NOC衡量模型通过共享行为子空间传递功能信息的强度。在多个模型家族中,这些图表高度压缩且部分可跨架构对齐。贡献空间图表展现出更强的架构鲁棒性共核,而激活空间图表保留更明显的家族特异性结构。通过受控后训练追踪图表变化,发现一致的不对称性:监督微调显著改变继承的行为几何,而奖励优化则在保持底层图表不变的前提下改变行为。这一几何视角为两类目标的机制差异提供了统一框架:SFT倾向于重写行为几何,而奖励优化主要重加权。

原文摘要 · Abstract (English)

Post-training can substantially alter language-model behavior, yet aggregate behavior rates do not reveal whether training removes an existing mechanism, creates a new one, or changes how an inherited mechanism is used. We study this question through two mechanistically distinct failures, repetition as a decoding-attractor pathology and sycophancy as a preference-related alignment failure. We introduce behavioral manifold analysis, which isolates behavior-specific geometry by selecting sparse behavior-associated coordinates and lifting them into low-dimensional local charts. We construct these charts in two complementary spaces. ACT captures runtime activation states, while NOC quantifies how strongly the model routes functional information flow through the shared behavior-associated subspace. Across multiple model families, the resulting charts are highly compressed and partially alignable across architectures. Contribution-space charts expose a more architecture-robust shared core, whereas activation-space charts retain stronger family-specific structure. Tracking these charts through controlled post-training reveals a consistent asymmetry. Supervised fine-tuning substantially alters the inherited behavioral geometry, whereas reward optimization changes behavior while largely preserving the underlying chart. This geometric perspective provides a unified framework for understanding the mechanistic distinction between the two objectives. SFT tends to rewrite behavioral geometry, whereas reward optimization primarily reweights it. Code is available at https://github.com/ronglingze/Manifold-Analysis

模型机制行为分析微调策略几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。