新基准GraphAllocBench让多目标强化学习适应用户偏好,更真实地评估算法性能。
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
- 基于城市资源分配的图结构环境,支持灵活定制目标与偏好条件。
- 提出PNDS和OS两项新指标,揭示超体积指标忽略的算法失效模式。
- 适合研究多目标强化学习、偏好建模及图神经网络在复杂任务中的应用。
多目标强化学习中的偏好条件策略学习(PCPL)通过将单一策略依用户偏好进行条件化,近似多样帕累托最优解,实现无需重训练即可运行时调整权衡。然而现有PCPL基准多局限于简单任务和固定环境,缺乏真实性和可扩展性。为此,我们提出GraphAllocBench,一个基于新型图结构资源分配沙盒CityPlannerEnv的灵活基准。该基准提供丰富问题集,支持自定义目标函数、变化偏好条件、复杂帕累托前沿及高维可扩展性。我们进一步提出两个补充指标:非支配解比例(PNDS)与排序得分(OS),用于衡量预测可靠性与偏好一致性,弥补广泛使用的超体积指标的不足。通过多个先进PCPL算法及我们提出的MLP与图感知的PCPL-PPO基线实验,结果显示GraphAllocBench揭示了超体积指标无法捕捉的显著失败模式,而新指标可有效识别;同时推动图神经网络等方法在复杂高维分配任务中的应用。用户可自由调整目标、偏好与分配规则,使GraphAllocBench成为推进PCPL研究的通用可扩展测试平台。
原文摘要 · Abstract (English)
Preference-Conditioned Policy Learning (PCPL) in Multi-Objective Reinforcement Learning (MORL) approximates diverse Pareto-optimal solutions by conditioning a single policy on user-specified preferences, enabling run-time adaptation to arbitrary trade-offs without retraining. However, existing PCPL benchmarks are largely restricted to toy tasks and fixed environments, limiting their realism and scalability. To address this gap, we introduce GraphAllocBench, a flexible benchmark built on CityPlannerEnv, a novel graph-based resource allocation sandbox inspired by city management. GraphAllocBench provides a rich suite of problems with customizable objective functions, varying preference conditions, complex Pareto Fronts, and high-dimensional scalability. We further propose two supplementary metrics -- Proportion of Non-Dominated Solutions (PNDS) and Ordering Score (OS) -- that capture prediction reliability and preference consistency while complementing the widely used hypervolume metric. Through experiments with several state-of-the-art PCPL algorithms and our own MLP and graph-aware PCPL-PPO baseline, we show that GraphAllocBench exposes distinct failure modes that hypervolume alone does not capture but our supplementary metrics reveal, while motivating graph-based approaches such as Graph Neural Networks (GNNs) for scaling to complex, high-dimensional allocation tasks. By letting users freely vary objectives, preferences, and allocation rules, GraphAllocBench serves as a versatile and extensible testbed for advancing PCPL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。