提出一种自适应学习率调度方法,平衡学习效率与努力成本。
Optimal Learning Rate Schedule for Balancing Effort and Performance
- 基于最优控制理论推导出闭式学习率公式,依赖当前与预期性能。
- 在模拟中复现数值优化的学习率调度,且跨任务和架构通用。
- 结合情景记忆机制实现生物可解释的近优学习行为,适合认知建模研究。
高效学习是生物体与人工智能体的核心挑战。为实现有效学习,智能体需调节学习速度,在快速进步与努力、不稳定性或资源消耗之间取得平衡。本文提出一个规范性框架,将该问题建模为最优控制过程:智能体在学习过程中最大化累积表现,同时承担学习成本。由此目标导出学习率的闭式解,其形式为仅依赖当前及未来预期表现的闭环控制器。在温和假设下,该解具有跨任务与模型架构的泛化能力,并在简单学习模型中复现了数值优化的学习率调度。我们进一步分析了智能体与任务参数如何通过开环控制影响学习率设计。由于最优策略依赖对未来的预期表现,该框架预测过度自信或低估会影响参与度与坚持性,将学习速度控制与自我调节学习理论联系起来。此外,我们证明一种简单的情景记忆机制可通过回忆相似过往经验来近似所需的表现预期,提供了一条生物合理的近优行为路径。这些结果共同构建了一个规范且生物可解释的学习速度控制框架,统一整合了自我调节学习、努力分配与情景记忆估计。
原文摘要 · Abstract (English)
Learning how to learn efficiently is a fundamental challenge for biological agents and a growing concern for artificial ones. To learn effectively, an agent must regulate its learning speed, balancing the benefits of rapid improvement against the costs of effort, instability, or resource use. We introduce a normative framework that formalizes this problem as an optimal control process in which the agent maximizes cumulative performance while incurring a cost of learning. From this objective, we derive a closed-form solution for the optimal learning rate, which has the form of a closed-loop controller that depends only on the agent's current and expected future performance. Under mild assumptions, this solution generalizes across tasks and architectures and reproduces numerically optimized schedules in simulations. In simple learning models, we can mathematically analyze how agent and task parameters shape learning-rate scheduling as an open-loop control solution. Because the optimal policy depends on expectations of future performance, the framework predicts how overconfidence or underconfidence influence engagement and persistence, linking the control of learning speed to theories of self-regulated learning. We further show how a simple episodic memory mechanism can approximate the required performance expectations by recalling similar past learning experiences, providing a biologically plausible route to near-optimal behaviour. Together, these results provide a normative and biologically plausible account of learning speed control, linking self-regulated learning, effort allocation, and episodic memory estimation within a unified and tractable mathematical framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。