arXiv:2605.06951cs.AIcs.LG2026-05

从多样专家行为中同时学习共性约束与个体偏好

Multi-Objective Constraint Inference using Inverse reinforcement learning

论文配图:Multi-Objective Constraint Inference using Inverse reinforcement learning
图 1 · 摘自论文原文
  • 基于逆强化学习,联合建模多目标差异与共享约束
  • 在网格世界任务中预测性能显著优于基线方法
  • 适合需要理解多元决策者意图的智能系统设计

约束推断在确保强化学习智能体符合安全边界和操作规范方面至关重要,通常通过观察专家示范实现。然而,现有方法多假设示范来源一致(即单一专家或目标相同的多位专家),难以捕捉个体偏好,且计算效率有限。本文提出多目标约束推断(MOCI)框架,能够从异质专家轨迹中联合提取共享约束与个体偏好,有效建模不同目标下的多样化甚至冲突行为。实证评估表明,MOCI在标准网格世界基准上显著优于现有基线,在预测性能上表现更优,同时保持了可接受的计算效率。结果证明MOCI是一种准确、灵活且计算可行的约束推断与偏好学习方法,适用于真实场景。

原文摘要 · Abstract (English)

Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observing expert demonstrations. However, existing approaches typically assume homogeneous demonstrations (i.e., generated by a single expert or multiple experts with identical objectives). They also have limited ability to capture individual preferences and often suffer from computational inefficiencies. In this paper, we introduce Multi-Objective Constraint Inference (MOCI), a novel framework designed to jointly extract shared constraints and individual preferences from heterogeneous expert trajectories, where multiple experts pursue different objectives. MOCI effectively models and learns from diverse, and potentially conflicting, behaviors. Empirical evaluations demonstrate that MOCI significantly outperforms existing baselines, achieving improved predictive performance, and maintaining competitive computational efficiency on a standard grid-world benchmark. These results establish MOCI as an accurate, flexible, and computationally practical approach for real-world constraint inference and preference learning tasks.

逆强化学习多目标学习偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。