arXiv:2510.02945cs.LGcs.AI2025-10

提出新型风险度量,让强化学习模型持续适应新环境时更稳健。

Ergodic Risk Measures: Towards a Risk-Aware Foundation for Continual Reinforcement Learning

  • 引入'遍历风险度量'新概念,适配持续学习场景
  • 证明传统风险度量在持续学习中不适用,新方法可避免遗忘
  • 适合关注模型长期稳定性与抗干扰能力的研究者

持续强化学习旨在形式化强化学习中的终身学习与无限适应。当前研究几乎均基于风险中性决策,即优化长期期望性能。本文首次从风险感知视角系统构建持续强化学习的理论基础,指出经典风险度量理论在持续学习中存在根本不兼容问题。基于此,我们提出一类新的遍历风险度量,并证明其与持续学习框架相容。最后通过案例研究与实证结果验证了该方法在持续风险感知学习中的直观有效性。

原文摘要 · Abstract (English)

Continual reinforcement learning (continual RL) seeks to formalize the notions of lifelong learning and endless adaptation in RL. In particular, the aim of continual RL is to develop RL agents that can maintain a careful balance between retaining useful information and adapting to new situations. To date, continual RL has been explored almost exclusively through the lens of risk-neutral decision-making, in which the agent aims to optimize the expected long-run performance. In this work, we present the first formal theoretical treatment of continual RL through the lens of risk-aware decision-making, in which the behaviour of the agent is directed towards optimizing a measure of long-run performance beyond the mean. In particular, we show that the classical theory of risk measures, widely used as a theoretical foundation in non-continual risk-aware RL, is, in its current form, incompatible with continual learning. Then, building on this insight, we extend risk measure theory into the continual setting by introducing a new class of ergodic risk measures, and showing that it is compatible with continual learning. Finally, we provide a case study of continual risk-aware learning, along with empirical results, which show the intuitive appeal of ergodic risk measures in continual settings.

持续学习风险感知强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。