发现神经网络持久记忆需稳定坐标系,否则会失效。
The Coordinate System Problem in Persistent Structural Memory for Neural Architectures
- 用挥发性路径网络构建稳定坐标系统,避免模型自学习坐标导致不稳定。
- 固定随机傅里叶特征可保坐标稳定,但需配合学习率调制才能实现有效迁移。
- 提出持久记忆两大要求:坐标稳定与平滑迁移机制,适用于结构化任务建模。
我们提出双视角信息素路径网络(DPPN),通过在潜在槽转移中沿持久信息素场路由稀疏注意力,揭示神经网络持久结构记忆的两个独立需求。基于最多每条件10个种子、5种模型变体和4个迁移目标的五轮渐进实验,发现持久记忆依赖于稳定坐标系,任何与模型共同学习的坐标系均内在不稳。识别出三大障碍:信息素饱和、表面结构纠缠、坐标不兼容;对比更新、多源蒸馏、匈牙利匹配、语义分解均无法在从零学习嵌入时解决不稳定性。固定随机傅里叶特征提供外在坐标,具备稳定、结构无关、信息丰富特性,但仅靠坐标稳定不足以实现迁移(10种子,p>0.05)。DPPN在任务内学习中优于Transformer和随机稀疏基线(AULC 0.700 vs 0.680 vs 0.670)。将路由偏置替换为学习率调制后消除负迁移:以预热信息素作为学习率先验,在同族任务上提升+0.003(17种子,p<0.05),且永不降低性能。基于外在坐标的结构补全函数带来+0.006同族任务增益,表明稳定与信息量之间的困境可通过学习函数部分破解。核心贡献为持久结构记忆的两项独立要求:(a) 坐标稳定性,(b) 平滑迁移机制。
原文摘要 · Abstract (English)
We introduce the Dual-View Pheromone Pathway Network (DPPN), an architecture that routes sparse attention through a persistent pheromone field over latent slot transitions, and use it to discover two independent requirements for persistent structural memory in neural networks. Through five progressively refined experiments using up to 10 seeds per condition across 5 model variants and 4 transfer targets, we identify a core principle: persistent memory requires a stable coordinate system, and any coordinate system learned jointly with the model is inherently unstable. We characterize three obstacles -- pheromone saturation, surface-structure entanglement, and coordinate incompatibility -- and show that neither contrastive updates, multi-source distillation, Hungarian alignment, nor semantic decomposition resolves the instability when embeddings are learned from scratch. Fixed random Fourier features provide extrinsic coordinates that are stable, structure-blind, and informative, but coordinate stability alone is insufficient: routing-bias pheromone does not transfer (10 seeds, p>0.05). DPPN outperforms transformer and random sparse baselines for within-task learning (AULC 0.700 vs 0.680 vs 0.670). Replacing routing bias with learning-rate modulation eliminates negative transfer: warm pheromone as a learning-rate prior achieves +0.003 on same-family tasks (17 seeds, p<0.05) while never reducing performance. A structure completion function over extrinsic coordinates produces +0.006 same-family bonus beyond regularization, showing the catch-22 between stability and informativeness is partially permeable to learned functions. The contribution is two independent requirements for persistent structural memory: (a) coordinate stability and (b) graceful transfer mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。