arXiv:2507.19647cs.ROcs.AI2025-07被引 6

用人类注视数据指导模仿学习,减少误判因果关系的问题。

GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning

  • 利用专家注视数据设计正则化损失,引导模型关注因果特征。
  • 在Atari和CARLA上比行为克隆提升179%和76%性能。
  • 提升模型可解释性,适合需要可靠决策的高风险场景。

模仿学习(IL)通过将任务视为监督学习来让智能体从人类专家演示中学习,但常因因果混淆而表现不佳——即错误地将虚假相关关系当作因果关系,导致在分布外环境中性能下降。为此,本文提出基于注视的正则化方法(GABRIL),利用数据收集阶段的人类注视数据,在模仿学习中引导表征学习。GABRIL通过正则化损失促使模型聚焦于专家注视所指示的因果相关特征,从而缓解混淆变量的影响。我们在Atari环境和CARLA中的Bench2Drive基准上构建了人类注视数据集,并应用该方法进行验证。实验结果表明,与行为克隆相比,GABRIL在Atari上的性能提升达179%,在CARLA设置下提升76%;此外,该方法相较传统模仿学习模型提供了额外的可解释性。

原文摘要 · Abstract (English)

Imitation Learning (IL) is a widely adopted approach which enables agents to learn from human expert demonstrations by framing the task as a supervised learning problem. However, IL often suffers from causal confusion, where agents misinterpret spurious correlations as causal relationships, leading to poor performance in testing environments with distribution shift. To address this issue, we introduce GAze-Based Regularization in Imitation Learning (GABRIL), a novel method that leverages the human gaze data gathered during the data collection phase to guide the representation learning in IL. GABRIL utilizes a regularization loss which encourages the model to focus on causally relevant features identified through expert gaze and consequently mitigates the effects of confounding variables. We validate our approach in Atari environments and the Bench2Drive benchmark in CARLA by collecting human gaze datasets and applying our method in both domains. Experimental results show that the improvement of GABRIL over behavior cloning is around 179% more than the same number for other baselines in the Atari and 76% in the CARLA setup. Finally, we show that our method provides extra explainability when compared to regular IL agents.

模仿学习注视追踪因果推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。