多目标强化学习需同时考虑不同时间尺度的非线性效用影响
It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning

- 提出跨时间尺度的非线性效用建模框架
- 证明现有方法在混合时序效应下存在理论缺陷
- 适合研究多目标决策与效用建模的学者
在处理多个奖励信号和非线性效用时,时间尺度至关重要。本文指出,当前主流的多目标强化学习方法(SER 和 ESR)以及后续特征(successor features)均不充分。尽管这些方法分别处理不同时间尺度上的非线性效用影响,但均未考虑在同一个决策问题中,不同时间尺度的效应可能同时发生。通过直观与数值示例,本文论证了这一现象的现实可能性,揭示了文献中的重要且非平凡的理论空白。
原文摘要 · Abstract (English)
Time is of the essence when dealing with multiple reward signals and non-linear utility. In this paper we argue that the current main approaches in multi-objective RL (SER and ESR), and successor features, are insufficient. While each approach deals with non-linear effects on user utility on different timescales, none of them take into account that different effects happening on different timescales can happen within the same decision problem. We motivate that this can indeed be the case by an example, both intuitively and numerically, leading to a new perspective, and a significant and non-trivial gap in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。