arXiv:2607.09842cs.LGcs.CL2026-07

发现多模态指令微调让模型用大小而非方向编码身份信息

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

论文配图:From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States
图 1 · 摘自论文原文
  • 通过曲率分布距离分析隐藏层轨迹,揭示身份提示的几何指纹变化
  • 指令微调后身份编码从方向转为幅度,基线模型方向差异显著(p=0.002)
  • 该现象仅在多模态指令微调中出现,适合研究模型表征机制的学者

我们研究了身份指定型系统提示是否在四种开放权重Transformer语言模型的隐藏状态轨迹中产生可区分的几何指纹,这些模型涵盖四种训练阶段:无训练(Gemma-4-E4B base)、多模态RLHF(Gemma-4-E4B-it)、RL蒸馏(DeepSeek-R1-Distill-Qwen-7B)和SFT(Qwen2.5-7B-Instruct)。比较三种提示条件(身份轴提示、长度匹配的通用助手提示、26令牌基线),使用五种几何度量,核心是基于k-NN轨迹图上边际分布的Ollivier-Ricci曲率的1-Wasserstein距离。分析基于轨迹级置换检验,包含多重几何控制(教师强制内容控制、时序链与k-NN拓扑对比、ABT投影k-NN、角度与欧氏图构建、5000次置换以验证边界统计量)。核心发现为:在指令微调边界处,身份编码发生定性重构——基线模型中指纹以方向编码为主(角度分离0.034,p=0.002,角度k-NN下);多模态指令微调模型中其迁移至幅度(角度分离降至p=0.439,欧氏仍显著,p=0.042,首个生成状态的平均范数反转长度排序,身份提示最低)。此方向到幅度的重构仅出现在多模态指令微调,其余训练方式中未见。教师强制控制表明,约30%的自由运行余弦信号来自提示驱动效应。我们提出将W_1应用于边际曲率分布作为独立方法贡献。

原文摘要 · Abstract (English)

We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state trajectories of four open-weight transformer language models spanning four post-training regimes: no training (Gemma-4-E4B base), multimodal RLHF (Gemma-4-E4B-it), RL distillation (DeepSeek-R1-Distill-Qwen-7B), and SFT (Qwen2.5-7B-Instruct). Three prompt conditions (an identity-specifying axis prompt, a length-matched generic-assistant prompt, and a 26-token vanilla baseline) are compared via five geometric metrics, principally the 1-Wasserstein distance between edge-wise distributions of Ollivier-Ricci curvature on k-NN trajectory graphs. Claims rest on trajectory-level permutation tests with multiple geometric controls (teacher-forced content controls, temporal-chain vs k-NN topology, ABT-projected k-NN, angular vs Euclidean graph construction, B=5000 permutations on borderline statistics). The central finding is a qualitative reorganization of identity encoding across the instruction-tuning boundary: in the base model the fingerprint is direction-coded (separation 0.034, p=0.002 under angular k-NN); in the multimodal instruction-tuned model it migrates into the magnitude (angular separation collapses to p=0.439 while Euclidean survives at p=0.042, and the mean norm of the first generated state inverts its length-ordering, being lowest for the identity prompt). This direction-to-magnitude reorganization is specific to the multimodal instruction-tuning regime, absent under RL distillation and SFT. A teacher-forced control attributes ~30% of the free-running cosine signal to prompt-driven effects. We position W_1 on edge-wise Ollivier-Ricci distributions on k-NN trajectory graphs as a methodological contribution of independent interest.

模型表征指令微调几何分析身份编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。