提出状态追踪的误差控制机制,揭示现有模型为何无法长期保持状态稳定。
Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
- 从误差控制角度重新审视循环模型的状态追踪能力
- 发现仿射网络无法纠正状态区分方向的误差,导致追踪失效
- 适合关注长时依赖、状态建模鲁棒性的研究者阅读
循环架构中的状态追踪理论长期聚焦于表达能力:固定结构能否实现一组符号转移规则。本文认为,误差控制同样关键,即隐藏状态在区分符号状态方向上的漂移动态。我们证明,包含状态空间模型和线性注意力的仿射循环网络一旦保留状态表示,便无法纠正状态分离子空间中的误差。因此,实际的仿射追踪器并未学习到鲁棒的状态追踪,而是学习了受累积状态相关误差约束的有限时域解。我们刻画了这一失败机制,指出追踪可读性仅在类内扩散相对于初始类间分离较小时成立。在群组状态追踪任务上,我们实证验证该崩溃可预测:当可区分性比率跨越训练解码器的可读阈值时,追踪即告失效。所有训练模型中,此交叉点准确预示下游准确率下降的时间窗口。结果表明,鲁棒状态追踪不仅取决于架构的理论表达能力,更关键的是其误差控制能力。
原文摘要 · Abstract (English)
The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretically realize a set of symbolic transition rules. We argue that equally important is error control, the dynamics governing hidden-state drift along the directions that distinguish symbolic states. We prove that affine recurrent networks, a class of models encompassing State-Space Models and Linear Attention, cannot correct errors along state-separating subspaces once they preserve state representations. Consequently, practical affine trackers do not learn robust state tracking; rather, they learn finite horizon solutions governed by accumulated state-relevant error. We characterize the mechanics of this failure, showing that tracking remains readable only while the accumulating within-class spread remains small relative to the initial between-class separation. We demonstrate empirically on group state-tracking tasks that this breakdown is predictable: tracking collapses when the distinguishability ratio crosses the readability threshold of the trained decoder. Across trained models, the point of this crossing predicts the horizon at which downstream accuracy fails. These results establish that robust state tracking is determined not only by an architecture's theoretical expressivity but crucially by its error control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。