arXiv:2603.21988cs.LGcs.AI2026-03中稿 · 4th World Conferen…

提出轨迹归因方法,解释多目标强化学习中不同行为对权衡结果的影响。

TREX: Trajectory Explanations for Multi-Objective Reinforcement Learning

  • 基于轨迹归因,将专家策略生成的轨迹按语义分段并聚类。
  • 通过移除特定行为段,量化其对帕累托权衡结果的影响,偏差达15%~30%。
  • 适用于需理解多目标决策过程的研究者或工业应用中的可解释性需求。

强化学习在多个领域展现出解决复杂决策问题的能力,通过与环境交互优化奖励信号。然而,许多现实场景涉及多个可能冲突的目标,难以用单一标量奖励表示。多目标强化学习(MORL)通过同时优化多个目标,显式处理它们之间的权衡。但强化学习模型的“黑箱”特性使得所选目标权衡背后的决策过程不清晰。现有可解释强化学习(XRL)方法通常针对单标量奖励设计,未考虑针对不同目标或用户偏好的解释。为此,本文提出TREX:一种基于轨迹归因的多目标强化学习可解释性框架。TREX从学习到的专家策略中生成轨迹,覆盖不同用户偏好,并将轨迹聚类为语义上连贯的时间片段。通过训练排除特定聚类的互补策略,测量其在观测奖励和动作上相对于原始专家策略的相对偏差,量化这些行为片段对帕累托权衡的影响。在多目标MuJoCo环境(半猫、蚂蚁、游动者)上的实验表明,该框架能有效识别并量化特定行为模式的作用。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has demonstrated its ability to solve complex decision-making problems in a variety of domains, by optimizing reward signals obtained through interaction with an environment. However, many real-world scenarios involve multiple, potentially conflicting objectives that cannot be easily represented by a single scalar reward. Multi-Objective Reinforcement Learning (MORL) addresses this limitation by enabling agents to optimize several objectives simultaneously, explicitly reasoning about trade-offs between them. However, the ``black box" nature of the RL models makes the decision process behind chosen objective trade-offs unclear. Current Explainable Reinforcement Learning (XRL) methods are typically designed for single scalar rewards and do not account for explanations with respect to distinct objectives or user preferences. To address this gap, in this paper we propose TREX, a Trajectory based Explainability framework to explain Multi-objective Reinforcement Learning policies, based on trajectory attribution. TREX generates trajectories directly from the learned expert policy, across different user preferences and clusters them into semantically meaningful temporal segments. We quantify the influence of these behavioural segments on the Pareto trade-off by training complementary policies that exclude specific clusters, measuring the resulting relative deviation on the observed rewards and actions compared to the original expert policy. Experiments on multi-objective MuJoCo environments - HalfCheetah, Ant and Swimmer, demonstrate the framework's ability to isolate and quantify the specific behavioural patterns.

强化学习可解释性多目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。