提出结构学习理论,解释系统如何在变化环境中识别并维护不同情境
Structural Decoupling: A Scaffold-Flow Theory of Generalization and Alignment
- 用'宽度'衡量完成任务所需最少情境数,定义结构学习核心
- 发现学习存在相变点,真实宽度可由收缩相似性算子估计
- 提出架构解耦原则,避免结构维持与内部优化互相干扰
非平稳多情境环境中的学习不仅需任务内泛化,还要求系统识别情境、正确路由输入、保留旧情境并随环境变化更新情境库。本文提出结构学习理论(StrLT)以填补这一结构性空白。StrLT 与 Vapnik 统计学习理论(SLT)互补:SLT 管控固定范式内的预测或控制(如漏斗);而 StrLT 管控结构范式的发现与维持(如陷阱)。核心对象为‘宽度’——覆盖问题所需的最小局部可行情境数。三个基本结果:宽度与 VC 维不可比较;学习在真实宽度处呈现相变;可通过收缩相似性(CS)算子估计宽度,将任务引发的非收缩性转化为谱分离。基于此,推导出结构解耦原则:维持结构框架的机制不应使用优化情境内流的同一梯度训练。这启发了架构解耦的梯度-流模型。最后指出,幻觉、奖励模型边界错误及欺骗对齐等安全问题,本质是结构解析或结构保持失败,而非单纯输出误差。
原文摘要 · Abstract (English)
Learning in non-stationary and multi-context environments requires more than ordinary within-task generalization. A system must also discover which contexts exist, route inputs to the correct context, preserve old contexts, and revise the context library when the environment changes. This paper presents Structural Learning Theory (StrLT) as a framework of filling this missing structural gap. StrLT complements Vapnik's Statistical Learning Theory (SLT): SLT governs the \emph{funnel}, prediction or control within a fixed regime; while StrLT governs the \emph{trap}, the discovery and maintenance of structural regimes. The core StrLT object is \emph{width}, the minimum number of locally feasible contexts needed to cover a problem. We summarize three basic results: width is incomparable with VC dimension; learning exhibits a phase transition at the true width; and width can be estimated by a contractive-similarity (CS) operator that converts task-induced non-contractivity into spectral separation. Under the StrLT framework, we explain how fixed-class structural learnability leads to a \emph{structural decoupling principle}: the mechanisms that maintain the structural scaffold should not be trained by the same gradients that optimize within-context flow. This principle motivates a scaffold-flow model in which alignment and generalization separate architecturally. Finally, we argue that several safety failures, including hallucination, reward-model boundary errors, and deceptive alignment, can be interpreted as scaffold-resolution or scaffold-preservation failures rather than merely output-level prediction errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。