arXiv:2601.00452cs.LG2026-01

用轨迹生成嵌入提升离线观察模仿学习的泛化能力

Imitation from Observations with Trajectory-Level Generative Embeddings

  • 构建时序扩散模型的潜在空间专家状态密度,生成平滑代理奖励
  • 在D4RL多任务基准上性能超越或持平现有方法
  • 适合处理稀疏专家数据和分布偏移严重的离线数据

我们研究离线观察模仿学习(LfO)问题,其中专家示范稀缺,可用的离线次优数据与专家行为分布相差甚远。现有分布匹配方法在此场景下表现不佳,因其施加严格支持约束并依赖脆弱的一步模型,难以从不完美数据中提取有效信号。为此,我们提出TGE:一种面向离线LfO的轨迹级生成嵌入方法,通过在基于离线轨迹数据训练的时序扩散模型的潜在空间中估计专家状态密度,构建稠密平滑的代理奖励。借助学习到的扩散嵌入的光滑几何结构,TGE能捕捉长时程时间动态,有效弥合不同支持集之间的差距,在离线数据与专家分布显著偏离时仍能提供稳健的学习信号。实验表明,该方法在D4RL多种运动与操作基准上一致优于或媲美现有离线LfO方法。

原文摘要 · Abstract (English)

We consider the offline imitation learning from observations (LfO) where the expert demonstrations are scarce and the available offline suboptimal data are far from the expert behavior. Many existing distribution-matching approaches struggle in this regime because they impose strict support constraints and rely on brittle one-step models, making it hard to extract useful signal from imperfect data. To tackle this challenge, we propose TGE, a trajectory-level generative embedding for offline LfO that constructs a dense, smooth surrogate reward by estimating expert state density in the latent space of a temporal diffusion model trained on offline trajectory data. By leveraging the smooth geometry of the learned diffusion embedding, TGE captures long-horizon temporal dynamics and effectively bridges the gap between disjoint supports, ensuring a robust learning signal even when offline data is distributionally distinct from the expert. Empirically, the proposed approach consistently matches or outperforms prior offline LfO methods across a range of D4RL locomotion and manipulation benchmarks.

模仿学习离线强化学习扩散模型轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。