arXiv:2604.27656cs.LGcs.AI2026-04

低维表示下模块化结构能有效避免干扰,高维则影响不大。

When Does Structure Matter in Continual Learning? Dimensionality Controls When Modularity Shapes Representational Geometry

论文配图:When Does Structure Matter in Continual Learning? Dimensionality Controls When Modularity Shapes Representational Geometry
图 1 · 摘自论文原文
  • 通过调节参数初始化控制表征维度,研究模块化与单网络的差异。
  • 低维时模块化网络实现任务子空间渐进对齐或正交,高维则无明显差异。
  • 适合关注持续学习中结构设计与表征几何关系的研究者。

为保留先前学习的表征,持续学习系统需在可塑性(获取新知识)与稳定性之间取得平衡。这一权衡影响表征跨任务复用:共享结构在任务相似时促进迁移,但可能引发新旧知识干扰。然而,结构性分离何时何地起作用仍不明确。本研究在受迁移-干扰研究启发的序列任务范式中,系统考察了网络架构、任务相似性与表征维度的共同作用。通过对比模块化循环网络与单模块基线,在不同任务相似度(低、中、高)和权重初始化尺度(影响学习模式)条件下,基于学习表征的有效维度进行实证分析。结果表明:在高维表征域中,架构影响微弱,因表征自由度足够容纳多任务而无显著干扰;而在低维(丰富)表征域中,架构至关重要——模块化网络呈现任务特定子空间的渐进式对齐(相似任务)、部分正交(中等差异)及强分离(差异大),而单模块网络缺乏此几何特性。这表明表征维度是决定结构性分离是否功能有效的关键变量,凸显自适应几何设计在持续学习系统中的核心作用。

原文摘要 · Abstract (English)

To preserve previously learned representations, continual learning systems must strike a balance between plasticity, the ability to acquire new knowledge, and stability. This stability-plasticity dilemma affects how representations can be reused across tasks: shared structure enables transfer when tasks are similar but may also induce interference when new learning disrupts existing representations. However, it remains unclear when and why structural separation influences this trade-off. In this study, we examine how network architecture, task similarity, and representational dimensionality jointly shape learning in a sequential task paradigm inspired by transfer-interference studies. We compare a task-partitioned modular recurrent network with a single-module baseline by systematically varying task similarity (low, medium, high) and the scale of weight initialization, which induces different learning regimes that we empirically characterize through the effective dimensionality of the learned representations. We find that architecture has minimal impact in high-dimensional regimes where representations are sufficiently unconstrained to accommodate multiple tasks without strong interference. In contrast, in lower-dimensional (rich) regimes, architectural separation is decisive: modular networks exhibit graded alignment of task-specific subspaces with overlap for similar tasks, partial orthogonalization for moderately dissimilar tasks, and stronger separation for dissimilar tasks. This graded geometry is absent in the single network baseline. Our findings suggest that representational dimensionality acts as a key organizing variable governing when structural separation becomes functionally relevant, and highlight adaptive geometry as a central principle for designing continual learning systems.

持续学习表征几何模块化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。