arXiv:2605.12000cs.LG2026-05

提出新方法,从多个专家数据中高效学习多目标最优策略。

Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation

论文配图:Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
图 1 · 摘自论文原文
  • 将冲突数据分组,相同行为数据合并,避免策略被劣化
  • 理论证明比独立处理每个专家数据更快收敛至最优
  • 在离散和连续任务上验证有效,适合复杂多目标场景

本文研究多目标模仿学习:在多目标马尔可夫决策过程(MOMDP)中,从多个帕累托最优专家的示范数据中恢复位于帕累托前沿的策略。标准模仿方法在此场景下表现不佳,因为简单聚合矛盾的专家轨迹可能导致被支配的策略。为此,我们提出多输出增强行为克隆(MA-BC),系统性地划分差异化的专家数据,同时对无行为冲突的状态-动作对进行聚合。理论上,我们证明了MA-BC以比独立处理每个专家数据集的任何学习器更快的统计速率收敛至帕累托最优策略。此外,我们建立了多目标模仿学习的新下界,证明了MA-BC是极小极大最优的。最后,我们在多种离散环境上实证验证了该算法,并基于理论洞察将其扩展至连续线性二次调节器(LQR)控制任务进行评估。

原文摘要 · Abstract (English)

This work investigates multi-objective imitation learning: the problem of recovering policies that lie on the Pareto front given demonstrations from multiple Pareto-optimal experts in a Multi-Objective Markov Decision Process (MOMDP). Standard imitation approaches are ill-equipped for this regime, as naively aggregating conflicting expert trajectories can result in dominated policies. To address this, we introduce Multi-Output Augmented Behavioral Cloning (MA-BC), an algorithm that systematically partitions divergent expert data while pooling state-action pairs where no behavior conflict is observed. Theoretically, we prove that MA-BC converges to Pareto-optimal policies at a faster statistical rate than any learner that considers each expert dataset independently. Furthermore, we establish a novel lower bound for multi-objective imitation learning, demonstrating that MA-BC is minimax optimal. Finally, we empirically validate our algorithm across diverse discrete environments and, guided by our theoretical insights, extend and evaluate MA-BC on a continuous Linear Quadratic Regulator (LQR) control task.

多目标学习模仿学习强化学习最优策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。