arXiv:2608.20442cs.LG2026-08

揭示隐性特质迁移的因果机制:优化器状态是关键载体。

Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer

  • 将参数与优化器动量视为统一训练状态,建立精确传输-估值关系。
  • 仅移植一阶动量即可引发后续参数与隐藏状态差异,且行为响应可恢复。
  • 该机制在多个模型与数据集上复现,适用于不同训练范式与架构。

隐性特质迁移使学生模型能从教师生成的数据中习得未语义表达的特征行为。现有研究解释了此类信号如何进入梯度,但未说明其如何在源数据移除后仍存在,或为何在后续训练中表现出不同符号。本文将参数与优化器动量视为单一训练状态,推导出一个精确的传输-估值恒等式,分离出独立于观察者的源扰动传播与未来延续赋予的价值。状态手术实验表明,一阶动量是因果载体:仅移植它,参数、隐藏状态和输出在切点处保持不变,但后续无源更新产生持续增长的参数与隐藏状态差异;若同时移植参数与一阶动量,则可恢复最终行为响应。向匹配的未来路径传递相同源诱导差异,产生负、近零、正三种Qwen效应(种子均值分别为 -0.658、+0.008 和 +0.658)。该排序在12个Llama-3.2-1B种子中经八次更新后重现,各路径的状态差范数几乎相等。当延续扩展至十六次更新时,所有配对种子的对比均增强。全时程代价状态可预测全部42个Qwen路径均值符号及全部21个已解决的Llama普通路径符号。独立于观察者的传输机制在Qwen、SmolLM2和Llama间复现;完整状态重复性还预测了非LoRA MNIST系统中的物理、隐藏与固定头响应,涵盖使用AdamW和动量SGD训练的CNN。综合结果揭示了隐性特质迁移的两阶段机制:优化器状态运输源扰动,后续训练决定其行为价值。

原文摘要 · Abstract (English)

Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically expressed. Recent work explains how such signals enter gradients, but not how they survive source removal or acquire different signs under later training. We treat parameters and optimizer moments as a single trainer state and derive an exact transport-valuation identity separating observer-independent propagation of the source perturbation from the value assigned by a future continuation and behavioral readout. State surgery identifies the first moment as a causal carrier. Transplanting it alone leaves parameters, hidden states, and outputs unchanged at the cut, yet source-free updates generate growing parameter and hidden-state differences; transplanting parameters with the first moment recovers the terminal behavioral response. Sending the same source-induced difference through matched futures produces negative, near-zero, and positive Qwen effects (-0.658, +0.008, and +0.658 seed means). This ordering recurs in all 12 Llama-3.2-1B seeds after eight updates, while state-difference norms remain nearly equal across routes. Both contrasts grow in every paired seed when the continuation extends to sixteen updates. A full-horizon costate predicts all 42 Qwen route-mean signs and all 21 resolved Llama ordinary-route signs. Observer-independent transport also replicates across Qwen, SmolLM2, and Llama, while the complete-state recurrence predicts physical, hidden, and fixed-head responses in non-LoRA MNIST systems, including CNNs trained with AdamW and momentum SGD. Together, these results identify a two-stage mechanism for subliminal trait transfer: optimizer state transports the source perturbation, and later training determines its behavioral value.

机器学习因果机制模型迁移优化器分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。