提出可逆行为学习,让模型调整后能精确恢复原状
On the Structural Limitations of Weight-Based Neural Adaptation and the Role of Reversible Behavioral Learning
- 将行为与参数解耦,实现行为的可逆卸载
- 实验显示可逆方法能数值精度回滚,传统方法有永久偏差
- 适合需要模型状态恢复的长期应用
神经网络通常通过共享参数进行微调、对齐训练和强化学习来适应。这些方法虽在短期优化中有效,但会导致模型基础行为发生长期改变。本文引入结构不可逆性概念,指出任务目标与模型表征身份纠缠,直接修改参数会使模型行为偏离原状态,且无法确定性还原,除非保存参数快照。为此提出可逆行为学习,使模型行为与身份参数解耦,可通过显式卸载过程确定性恢复。引入可恢复性因子作为行为可恢复性的归一化度量,并提供基于模型偏差的额外诊断工具。实验表明,可逆适应可在数值精度内实现回滚,而共享参数修改则存在持续的重置后偏差。
原文摘要 · Abstract (English)
Neural models are usually adapted through changes in parameters shared among model components via fine-tuning, alignment-based training, and reinforcement learning. These changes have been found effective in short-term optimization. However, they result in long-term alterations in the model's base behavior. In this study, we introduce the concept of structural irreversibility as a characteristic of shared-parameter model adaptation. This concept refers to the intertwining of task-specific objectives with the representational identity of the model. We show that when parameters are directly mutated, the resulting model behaves divergently from the original model. This divergence cannot be reversed deterministically without an explicit parameter snapshot. We introduce reversible behavioral learning, in which model behaviors are structurally dissociated from identity parameters and can be deterministically unloaded through an explicit unload process. We also introduce the Recoverability Factor as a normalized measure of behavioral recoverability and provide additional diagnostics based on model divergence. Experiments show that reversible model adaptation achieves rollback within numerical precision, whereas shared-parameter mutation exhibits persistent post-reset divergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。