arXiv:2411.04784cs.AIcs.LG2024-11被引 4

为多目标强化学习的解集提供聚类分析,帮助决策者快速发现策略规律。

Navigating Trade-offs: Policy Summarization for Multi-Objective Reinforcement Learning

  • 结合策略行为与目标值进行聚类,揭示策略与目标空间的关系。
  • 在4个环境中优于传统k-medoids方法,能有效识别解集中的模式。
  • 适合需要理解多目标权衡的决策者或系统设计者使用。

多目标强化学习(MORL)用于解决涉及多个目标的问题,其训练结果生成一组策略,每组策略在不同目标间呈现特定权衡(预期回报)。相比单一策略,MORL通过细粒度比较策略间的权衡提升了可解释性。然而,解集通常规模大且高维,每个策略(如神经网络)仅以目标值表示。本文提出一种基于策略行为和目标值的聚类方法,揭示策略行为与目标空间区域的关系。该方法使决策者能整体把握解集中的趋势与洞察,无需逐个分析每个策略。我们在四个多目标环境测试该方法,结果优于传统k-medoids聚类。此外,还通过案例研究展示了其在实际场景中的应用价值。

原文摘要 · Abstract (English)

Multi-objective reinforcement learning (MORL) is used to solve problems involving multiple objectives. An MORL agent must make decisions based on the diverse signals provided by distinct reward functions. Training an MORL agent yields a set of solutions (policies), each presenting distinct trade-offs among the objectives (expected returns). MORL enhances explainability by enabling fine-grained comparisons of policies in the solution set based on their trade-offs as opposed to having a single policy. However, the solution set is typically large and multi-dimensional, where each policy (e.g., a neural network) is represented by its objective values. We propose an approach for clustering the solution set generated by MORL. By considering both policy behavior and objective values, our clustering method can reveal the relationship between policy behaviors and regions in the objective space. This approach can enable decision makers (DMs) to identify overarching trends and insights in the solution set rather than examining each policy individually. We tested our method in four multi-objective environments and found it outperformed traditional k-medoids clustering. Additionally, we include a case study that demonstrates its real-world application.

强化学习多目标优化策略分析聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。