无标签视频中提取与真实动作强相关的潜在表示,提升机器人学习性能。
Boosting Action-Information via a Variational Bottleneck on Unlabelled Robot Videos
- 用变分信息瓶颈最大化潜在动作与真实动作的互信息
- 在无标签视频上实现动作相关表征,提升控制效果
- 适合做无需人工标注动作数据的机器人学习研究者
从示范中学习(LfD)通常依赖大量带动作标签的专家轨迹,严重限制了可用训练数据规模。一种有前景的替代方案是直接从无标签视频示范中学习。然而,我们发现现有方法倾向于编码与真实机器人动作关联性弱的潜在动作,导致控制性能不佳。为解决此问题,我们提出一种新框架,在无动作标签情况下显式最大化潜在动作与真实动作之间的互信息。该方法利用变分信息瓶颈提取与动作相关的信息,同时丢弃任务无关内容。我们提供了理论分析,证明该目标确实能最大化潜在动作与真实动作间的互信息。通过大量实验验证:在模拟机器人环境和真实机器人平台上,结果均表明该方法显著提升了互信息,并持续改善策略性能。
原文摘要 · Abstract (English)
Learning from demonstrations (LfD) typically relies on large amounts of action-labeled expert trajectories, which fundamentally constrains the scale of available training data. A promising alternative is to learn directly from unlabeled video demonstrations. However, we find that existing methods tend to encode latent actions that share little mutual information with the true robot actions, leading to suboptimal control performance. To address this limitation, we introduce a novel framework that explicitly maximizes the mutual information between latent actions and true actions, even in the absence of action labels. Our method leverage the variational information-bottleneck to extract action-relevant representations while discarding task-irrelevant information. We provide a theoretical analysis showing that our objective indeed maximizes the mutual information between latent and true actions. Finally, we validate our approach through extensive experiments: first in simulated robotic environments and then on real-world robotic platforms, the experimental results demonstrate that our method significantly enhances mutual information and consistently improves policy performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。