arXiv:2511.16476cs.LG2025-11

对比分析多目标强化学习中标量化方法的局限性,发现其易丢失解且难覆盖完整最优前沿。

Limitations of Scalarisation in MORL: A Comparative Study in Discrete Environments

  • 采用外层多策略评估法对比线性与切比雪夫标量化算法性能
  • 标量化方法在不同环境和最优前沿形状下表现差异大,常遗漏学习到的解
  • 内层多策略算法更稳健,适合动态不确定环境中的智能决策

标量化函数广泛用于多目标强化学习(MORL)算法以实现智能决策,但在复杂不确定环境中难以准确逼近帕累托前沿。本研究在离散动作与观测空间的MORL环境中,评估了多种经典算法。通过外层多策略方法,对比了基于线性标量化和切比雪夫标量化实现的单策略算法MO Q-Learning的性能;同时探讨了开创性的内层多策略算法Pareto Q-Learning。结果表明,标量化方法的表现高度依赖环境及帕累托前沿形状,常无法保留学习过程中发现的解,偏好特定解空间区域。此外,为全面采样帕累托前沿而寻找合适权重配置极为困难,限制了其在不确定场景的应用。相比之下,内层多策略算法展现出更强的鲁棒性与可扩展性,可能更适合动态、不确定环境中的智能决策。

原文摘要 · Abstract (English)

Scalarisation functions are widely employed in MORL algorithms to enable intelligent decision-making. However, these functions often struggle to approximate the Pareto front accurately, rendering them unideal in complex, uncertain environments. This study examines selected Multi-Objective Reinforcement Learning (MORL) algorithms across MORL environments with discrete action and observation spaces. We aim to investigate further the limitations associated with scalarisation approaches for decision-making in multi-objective settings. Specifically, we use an outer-loop multi-policy methodology to assess the performance of a seminal single-policy MORL algorithm, MO Q-Learning implemented with linear scalarisation and Chebyshev scalarisation functions. In addition, we explore a pioneering inner-loop multi-policy algorithm, Pareto Q-Learning, which offers a more robust alternative. Our findings reveal that the performance of the scalarisation functions is highly dependent on the environment and the shape of the Pareto front. These functions often fail to retain the solutions uncovered during learning and favour finding solutions in certain regions of the solution space. Moreover, finding the appropriate weight configurations to sample the entire Pareto front is complex, limiting their applicability in uncertain settings. In contrast, inner-loop multi-policy algorithms may provide a more sustainable and generalizable approach and potentially facilitate intelligent decision-making in dynamic and uncertain environments.

多目标强化学习标量化帕累托前沿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。