arXiv:2502.02867cs.CVcs.AI2025-02

让智能体跨域模仿专家动作,从视觉中提取不变特征。

Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations

  • 逐帧提取跨域不变特征,再组装成序列
  • 通过时间标签划分行为阶段,提升模仿精度
  • 适合视觉观测复杂、数据不完整的跨域场景

模仿学习(IL)使智能体无需奖励信号即可模仿专家行为,但在高维、噪声大且不完整的视觉观测下存在跨域挑战。为此,我们提出一种新方法——针对模仿学习的域不变逐帧特征提取(DIFF-IL),从单帧中提取域不变特征,并将其转化为序列以分离并复现专家行为。同时引入逐帧时间标记技术,按时间步分割专家行为并分配与上下文一致的奖励,从而提升任务表现。在多种视觉环境下的实验表明,DIFF-IL在处理复杂视觉任务时具有显著有效性。

原文摘要 · Abstract (English)

Imitation learning (IL) enables agents to mimic expert behavior without reward signals but faces challenges in cross-domain scenarios with high-dimensional, noisy, and incomplete visual observations. To address this, we propose Domain-Invariant Per-Frame Feature Extraction for Imitation Learning (DIFF-IL), a novel IL method that extracts domain-invariant features from individual frames and adapts them into sequences to isolate and replicate expert behaviors. We also introduce a frame-wise time labeling technique to segment expert behaviors by timesteps and assign rewards aligned with temporal contexts, enhancing task performance. Experiments across diverse visual environments demonstrate the effectiveness of DIFF-IL in addressing complex visual tasks.

模仿学习跨域泛化视觉感知特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。