提出新方法解决长序列非阿贝尔状态追踪难题,效果远超基线。
A Held-Out Transition-Pair Falsifier for Long-Horizon Non-Abelian State Tracking
- 设计保留对偶转移对的验证机制,阻止模型直接记忆局部转移
- 在 $S_3 \times S_3$ 基准上实现长达 1048576 步的零误差预测
- 适合研究长程依赖与非交换状态建模的学者参考
状态追踪揭示了序列模型的一个关键局限:相关信号并非观测标记的摘要,而是一个通过非交换变换演化的有序潜在状态。本文提出一种保留对偶转移对的伪造检测器,用于有限非阿贝尔群的状态追踪。训练阶段禁止特定有序生成对,评估阶段要求相同的局部模式,从而阻断直接的局部转移记忆路径。在控制性 $S_3 \times S_3$ 基准上,仅在长度为 8 的序列上训练的投影递归状态模型,在长达 1,048,576 个标记的评估序列上,五个随机种子下均实现 250/250 的零误差最终状态预测。匹配的原生读出基线(包括袋模型、GRU 和单配置结构化状态空间模型)在相同协议下仍接近随机水平。配备类似有限群原型读出的投影匹配 GRU、结构化 SSM 和袋模型基线同样表现接近随机。机制诊断显示,硬投影对应低同态误差、低状态一致性漂移和非平凡的换位子分离;而软投影则导致最终状态准确率下降。干净划分审计确认训练与评估分区间无重复字串重叠和无结构模板重叠。证据限于该受控有限群伪造器框架,而非一般架构排名。在此范围内,显式投影的非交换状态组合可作为长时程隐状态追踪的有效归纳偏置。
原文摘要 · Abstract (English)
State tracking exposes a sharp limitation of sequence models: the relevant signal is often not a summary of observed tokens, but an ordered latent state that evolves through non-commutative transformations. We introduce a held-out transition-pair falsifier for finite non-Abelian group tracking. The protocol forbids selected ordered generator pairs during training and requires the same local patterns during evaluation, blocking one direct local-transition memorization pathway. In a controlled $S_3 \times S_3$ benchmark, a projected recurrent state model trained only on length-8 sequences produces error-free final-state predictions (perfect 250/250 per horizon) through evaluation horizons up to 1,048,576 tokens across five seeds. Matched native-readout baselines, including bag, GRU, and a single-configuration structured state-space model, remain near floor under the same protocol. Projection-matched GRU, structured SSM, and bag baselines equipped with analogous finite-group prototype readouts also remain near chance under the same split. Mechanism diagnostics show that hard projection coincides with low homomorphism error, low state-consistency drift, and non-trivial commutator separation, while softened projection collapses final-state accuracy. Clean-split audits verify zero verbatim reduced-word overlap and zero structural-template overlap between training and evaluation partitions. The evidence is scoped to this controlled finite-group falsifier rather than to a general architecture ranking. Within that regime, explicit projected non-commutative state composition acts as a useful inductive bias for long-horizon hidden-state tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。