arXiv:2602.14252cs.AIcs.LG2026-02中稿 · publication at AAM…被引 2

用模仿学习让AI更准识别人类目标,尤其在行为不完美时仍可靠。

GRAIL: Goal Recognition Alignment through Imitation Learning

  • 通过模仿学习直接从演示轨迹中学习每个候选目标的策略
  • 在有系统偏差时F1分数提升超0.5,噪声轨迹下最高提升0.4
  • 适合需要理解不完美人类行为的AI对齐场景

从行为中理解智能体的目标是使人工智能系统与人类意图对齐的基础。现有目标识别方法通常依赖于最优的目标导向策略表示,但该表示可能与行为者的真实行为不符,从而影响目标的准确识别。为此,本文提出基于模仿学习的目标识别对齐方法(GRAIL),利用模仿学习和逆强化学习,直接从(可能非最优的)演示轨迹中为每个候选目标学习一个目标导向策略。通过单次前向传播对每个学习到的目标导向策略评分观测到的部分轨迹,GRAIL保持了经典目标识别的一次性推理能力,同时利用可捕捉非最优和系统性偏差行为的策略。在所评估领域中,当存在系统性偏差的最优行为时,GRAIL的F1分数提升超过0.5;在非最优行为下,提升约0.1-0.3;在噪声最优轨迹下,最高提升达0.4,且在完全最优设置下仍具竞争力。本工作推动了在不确定环境中解释智能体目标的可扩展、鲁棒模型发展。

原文摘要 · Abstract (English)

Understanding an agent's goals from its behavior is fundamental to aligning AI systems with human intentions. Existing goal recognition methods typically rely on an optimal goal-oriented policy representation, which may differ from the actor's true behavior and hinder the accurate recognition of their goal. To address this gap, this paper introduces Goal Recognition Alignment through Imitation Learning (GRAIL), which leverages imitation learning and inverse reinforcement learning to learn one goal-directed policy for each candidate goal directly from (potentially suboptimal) demonstration trajectories. By scoring an observed partial trajectory with each learned goal-directed policy in a single forward pass, GRAIL retains the one-shot inference capability of classical goal recognition while leveraging learned policies that can capture suboptimal and systematically biased behavior. Across the evaluated domains, GRAIL increases the F1-score by more than 0.5 under systematically biased optimal behavior, achieves gains of approximately 0.1-0.3 under suboptimal behavior, and yields improvements of up to 0.4 under noisy optimal trajectories, while remaining competitive in fully optimal settings. This work contributes toward scalable and robust models for interpreting agent goals in uncertain environments.

目标识别模仿学习行为理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。