提出评估多目标强化学习中偏好控制能力的新指标。
Controllability in preference-conditioned multi-objective reinforcement learning

- 设计新度量标准检测偏好输入是否可靠改变智能体行为。
- 发现主流指标无法识别偏好不敏感的智能体,存在评估漏洞。
- 适合关注人机对齐与可解释强化学习的研究者参考。
多目标强化学习(MORL)允许用户通过目标相对重要性表达偏好,但标准度量无法判断偏好变化是否可靠地引发预期行为改变,这一特性称为可控性。因此,偏好条件化智能体可能在标准MORL指标上表现良好,却对偏好输入不敏感。若无法可靠评估可控性,MORL提供的用户意图与智能体行为之间的符号接口便失效。主流MORL度量无法衡量偏好条件化智能体的可控性,这促使我们设计一个专门为此目的而生的补充度量。我们希望这些结果能推动社区讨论现有评估协议,以巩固多目标强化学习中偏好适应的进展,应用于更大更复杂的问题。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) allows a user to express preference over outcomes in terms of the relative importance of the objectives, but standard metrics cannot capture whether changes in preference reliably change the agent's behavior in the intended way, a property termed controllability. As a result, preference-conditioned agents can score well on standard MORL metrics while being insensitive to the preference input. If the ability to control agents cannot be reliably assessed, the symbolic interface that MORL provides between user intent and agent behavior is broken. Mainstream MORL metrics alone fail to measure the controllability of preference-conditioned agents, motivating a complementary metric specifically designed to that end. We hope the results spur discussion in the community on existing evaluation protocols to consolidate advances in preference adaptation in MORL to larger and more complex problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。