arXiv:2606.17889cs.LGcs.AI2026-06中稿 · the 2nd Workshop o…

低维表示下模块化架构能更好应对持续学习中的干扰问题。

Dimensionality Controls When Modularity Helps in Continual Learning

论文配图:Dimensionality Controls When Modularity Helps in Continual Learning
图 1 · 摘自论文原文
  • 通过调节权重尺度控制表征维度,对比模块化与单网络结构。
  • 低维时模块化网络形成任务相关的子空间,相似任务共享更多表征。
  • 揭示了表征维度是决定模块化是否有效的关键因素,适合研究持续学习机制的人参考。

组合式学习系统需在可塑性(获取新知识)与稳定性(保留旧知识)间取得平衡,尤其当任务共享结构且存在干扰风险时。本文在顺序A-B-A范式下,研究模块化架构、任务相似性与表征维度如何共同影响组合式持续学习,对比了任务分割的循环网络与单网络基线,并通过权重尺度调控实现高维与低维表征情形。在高维“懒惰”模式下,两种架构表现相近,内部几何结构也相似,表明此时显式模块化作用有限。而在低维“丰富”模式下,模块化网络发展出渐进式的任务特定子空间:相似任务重叠较多,中度不同时部分对齐,差异大时则分离,其组织方式比单网络更组合化且可解释。结果表明,初始化尺度所诱导的表征区间(与表征维度共变)是决定模块化在持续学习中是否具有功能优势的关键因素,支持将安全性和鲁棒性视为表征子空间的自适应分配问题,而非固定分离或共享。

原文摘要 · Abstract (English)

Compositional learning systems must balance plasticity, the ability to acquire new knowledge, with stability, the preservation of previously learned components, especially when tasks share structure and risk interference. We study how modular architecture, task similarity, and representational dimensionality jointly shape compositional continual learning in a sequential A-B-A paradigm, comparing a task-partitioned recurrent network to a single-network baseline while inducing high- and low-dimensional regimes via weight-scale manipulations. In a high-dimensional "lazy" regime, both architectures achieve similar performance and internal geometry, suggesting that explicit modular structure has little impact when representations are weakly constrained. In a lower-dimensional "rich" regime, modularity becomes decisive: the modular network develops graded task-specific subspaces that overlap for similar tasks, partially align for moderately dissimilar tasks, and separate for dissimilar tasks, yielding a more compositional and interpretable organization than the single network. These findings identify the representational regime induced by initialization scale, which co-varies with representational dimensionality, as a key factor governing when compositional, modular structure is functionally beneficial in continual learning, and support viewing safety and robustness as problems of adaptive allocation of representational subspaces rather than fixed separation versus sharing.

持续学习模块化表征维度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。