arXiv:2606.21321cs.LG2026-06

提出诊断方法,揭示多目标强化学习中隐藏的行为差异。

Objective-Behavior Alignment: Diagnostics for MORL Policy Selection

论文配图:Objective-Behavior Alignment: Diagnostics for MORL Policy Selection
图 1 · 摘自论文原文
  • 通过自动分析帕累托前沿,发现仅凭收益值无法识别的行为差异。
  • 在网格与连续控制任务中验证,复杂度升高仍有效。
  • 适合需理解策略行为差异的决策者或算法调试者。

现实世界决策常需同时优化多个相互竞争的目标。强化学习中通常通过标量函数将奖励信号合并为单一目标,但该方法脆弱:权重微小变化即可引发截然不同的策略。多目标强化学习(MORL)则生成一组显式体现目标权衡的策略。然而,这些策略通常仅通过其价值向量呈现,可能掩盖显著的行为差异:不同轨迹的策略在仅以期望回报评估时可能看似无异。本文提出一种探索性诊断流程,可自动揭示仅凭目标值无法反映的帕累托前沿上的行为变异,提供定量与可视化工具支持策略检查。我们在简单网格示例上验证该方法,并扩展至连续控制基准,证明其在问题复杂度提升时仍具有效性。

原文摘要 · Abstract (English)

Real-world decision-making often requires optimizing multiple competing objectives simultaneously. In reinforcement learning (RL), this is typically addressed by combining reward signals into a single scalar objective via a scalarization function, which can be fragile: small changes in the weights can induce drastically different policies. Multi-objective reinforcement learning (MORL) instead produces sets of policies that explicitly represent trade-offs between objectives. However, these policies are typically presented to the decision maker only through their value vectors, which can obscure substantial behavioral variation: policies that induce distinct trajectories may appear indistinguishable when evaluated solely by expected returns. We propose an exploratory diagnostic workflow that automatically highlights behavioral variation along the Pareto front that objective values alone do not reveal, providing both quantitative and visual tools to support policy inspection. We validate our approach on simple grid examples and scale it to continuous control benchmarks, demonstrating that it remains effective as problem complexity increases.

多目标强化学习策略诊断行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。