从行为示范中推断多目标决策者的偏好,无需用户交互
Inferring Preferences from Demonstrations in Multi-objective Reinforcement Learning
- 基于动态权重构建偏好推理算法,从示范中学习目标偏好
- 在3个环境上表现优于基线,精度与效率均显著提升
- 适用于次优示范,无需额外交互,适合真实场景应用
许多决策问题涉及多个目标,而人类或智能体对各目标的偏好往往难以直接获取。然而,决策者的行为示范通常可得。本文提出一种动态权重偏好推理(DWPI)算法,能够从示范中推断多目标决策中的偏好。该算法在三个多目标马尔可夫决策过程(Deep Sea Treasure、Traffic、Item Gathering)上进行了评估,并与两种现有算法对比。实验结果表明,相比基线算法,DWPI在推理准确性和时间效率上均有显著提升。该算法在处理次优示范时仍保持稳定性能,且推理过程无需用户交互,仅需示范数据。本文提供了算法正确性证明与复杂度分析,并在不同示范表示下进行了统计性能评估。
原文摘要 · Abstract (English)
Many decision-making problems feature multiple objectives where it is not always possible to know the preferences of a human or agent decision-maker for different objectives. However, demonstrated behaviors from the decision-maker are often available. This research proposes a dynamic weight-based preference inference (DWPI) algorithm that can infer the preferences of agents acting in multi-objective decision-making problems from demonstrations. The proposed algorithm is evaluated on three multi-objective Markov decision processes: Deep Sea Treasure, Traffic, and Item Gathering, and is compared to two existing preference inference algorithms. Empirical results demonstrate significant improvements compared to the baseline algorithms, in terms of both time efficiency and inference accuracy. The DWPI algorithm maintains its performance when inferring preferences for sub-optimal demonstrations. Moreover, the DWPI algorithm does not necessitate any interactions with the user during inference - only demonstrations are required. We provide a correctness proof and complexity analysis of the algorithm and statistically evaluate the performance under different representation of demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。