让模型像大脑一样自我进化,动态调整结构适应长期学习。
Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- 模型能自主调整层级数量、结构和更新频率,模仿神经可塑性。
- 在语言建模和长上下文推理中表现优于现有方法,支持持续学习。
- 适合研究通用人工智能与长期适应性系统的学者和工程师。
当前机器学习模型(包括大语言模型)在静态任务中表现优异,但在非平稳环境中因架构僵化而难以持续适应与长期学习。本文提出动态嵌套层级结构,作为嵌套学习范式的演进,使模型能在训练或推理过程中自主调整优化层级数、嵌套结构和更新频率,受神经可塑性启发,实现无需预设约束的自演化。该机制克服了现有模型的顺行性遗忘问题,通过动态压缩上下文流并适应分布漂移,实现真正意义上的终身学习。基于严格的数学建模、收敛性证明、表达能力边界分析及不同场景下的亚线性遗憾理论,结合语言建模、持续学习和长上下文推理的实证验证,动态嵌套层级结构为构建自适应、通用智能奠定了基础。
原文摘要 · Abstract (English)
Contemporary machine learning models, including large language models, exhibit remarkable capabilities in static tasks yet falter in non-stationary environments due to rigid architectures that hinder continual adaptation and lifelong learning. Building upon the nested learning paradigm, which decomposes models into multi-level optimization problems with fixed update frequencies, this work proposes dynamic nested hierarchies as the next evolutionary step in advancing artificial intelligence and machine learning. Dynamic nested hierarchies empower models to autonomously adjust the number of optimization levels, their nesting structures, and update frequencies during training or inference, inspired by neuroplasticity to enable self-evolution without predefined constraints. This innovation addresses the anterograde amnesia in existing models, facilitating true lifelong learning by dynamically compressing context flows and adapting to distribution shifts. Through rigorous mathematical formulations, theoretical proofs of convergence, expressivity bounds, and sublinear regret in varying regimes, alongside empirical demonstrations of superior performance in language modeling, continual learning, and long-context reasoning, dynamic nested hierarchies establish a foundational advancement toward adaptive, general-purpose intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。