用物理与控制理论推导出最优持续学习策略,缓解灾难性遗忘。
Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
- 结合统计物理与最优控制,推导出任务学习顺序的理论最优解。
- 理论预测在相似任务间交替可显著减少遗忘,实验证实有效。
- 适合研究持续学习机制或设计高效训练协议的研究者。
人工神经网络在顺序学习多个任务时常遭遇灾难性遗忘,即新任务训练会损害旧任务性能。尽管已有理论工作通过合成框架分析学习曲线,但现有训练协议多依赖启发式方法,缺乏理论支撑以评估其最优性。本文结合统计物理推导的精确训练动态方程与最优控制方法,应用于教师-学生模型中的持续学习与多任务问题,构建了最大化性能同时最小化遗忘的任务选择理论。理论分析揭示了非平凡但可解释的缓解遗忘策略,阐明了任务相似性对遗忘的影响机制。最后,我们在真实数据上验证了理论结果的有效性。
原文摘要 · Abstract (English)
Artificial neural networks often struggle with catastrophic forgetting when learning multiple tasks sequentially, as training on new tasks degrades the performance on previously learned tasks. Recent theoretical work has addressed this issue by analysing learning curves in synthetic frameworks under predefined training protocols. However, these protocols relied on heuristics and lacked a solid theoretical foundation assessing their optimality. In this paper, we fill this gap by combining exact equations for training dynamics, derived using statistical physics techniques, with optimal control methods. We apply this approach to teacher-student models for continual learning and multi-task problems, obtaining a theory for task-selection protocols maximising performance while minimising forgetting. Our theoretical analysis offers non-trivial yet interpretable strategies for mitigating catastrophic forgetting, shedding light on how optimal learning protocols modulate established effects, such as the influence of task similarity on forgetting. Finally, we validate our theoretical findings with experiments on real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。