arXiv:2602.04131cs.LG2026-02

提出可灵活调整时间与风险偏好的强化学习框架

Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting

  • 用可变折扣函数替代固定指数折扣,解耦时间与风险偏好
  • 多时域扩展修复现有方法的优化缺陷,提升策略稳定性
  • 适合安全关键场景,如自动驾驶、医疗决策等风险敏感应用

分布式强化学习(Distributional RL)在安全关键领域日益受到重视,因其能优化风险敏感目标。然而,折扣因子常被当作固定参数或可调超参,未被充分关注其对策略的影响。文献表明,折扣函数决定了智能体的时间偏好,而指数折扣无法完全刻画这一特性。为此,本文提出一种支持灵活未来奖励折扣并优化风险度量的分布式强化学习新框架。我们从理论上分析算法最优性,证明多时域扩展解决了现有方法的缺陷,并通过大量实验验证了方法的鲁棒性。结果表明,折扣是决策问题中捕捉更丰富的时间与风险偏好特征的核心,对真实世界安全关键应用具有潜在意义。

原文摘要 · Abstract (English)

Distributional reinforcement learning (RL) is a powerful framework increasingly adopted in safety-critical domains for its ability to optimize risk-sensitive objectives. However, the role of the discount factor is often overlooked, as it is typically treated as a fixed parameter of the Markov decision process or tunable hyperparameter, with little consideration of its effect on the learned policy. In the literature, it is well-known that the discounting function plays a major role in characterizing time preferences of an agent, which an exponential discount factor cannot fully capture. Building on this insight, we propose a novel framework that supports flexible discounting of future rewards and optimization of risk measures in distributional RL. We provide a technical analysis of the optimality of our algorithms, show that our multi-horizon extension fixes issues raised with existing methodologies, and validate the robustness of our methods through extensive experiments. Our results highlight that discounting is a cornerstone in decision-making problems for capturing more expressive temporal and risk preferences profiles, with potential implications for real-world safety-critical applications.

强化学习风险敏感折扣函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。