通过多视角重建学习更贴近真实动作的潜在动作表示。
MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
- 利用多视角视频间重建约束,让潜在动作捕捉跨视角共性。
- 在Bridge V2上与真实动作互信息更高,分布外泛化能力更强。
- 适合需要强动作语义表征的视觉-语言-动作预训练任务。
从多样化人体视频中学习的潜在动作可作为视觉-语言-动作(VLA)预训练的伪标签,但仅当其仍包含底层真实动作信息时才能提供有效监督。为此,我们提出多视角潜在动作模型(MVP-LAM),通过跨视角重建目标学习对真实动作高度敏感的潜在动作表示。该方法要求某一视角的潜在动作能解释另一视角的未来状态,从而减少对视角特异性线索的依赖。在Bridge V2数据集上,MVP-LAM生成的潜在动作更具动作中心性,与真实动作的互信息更高,且在分布外评估下仍保持优异的动作预测性能。将此潜在动作用于VLA预训练后,显著提升了下游操作任务在多个基准上的表现。代码与训练权重已公开于https://jmsnu.github.io。
原文摘要 · Abstract (English)
Latent actions learned from diverse human videos serve as pseudo-labels for vision-language-action (VLA) pretraining, but provide effective supervision only if they remain informative about the underlying ground-truth actions. For effective supervision, latent actions should contain information about the underlying actions even though they are inaccessible. We propose Multi-ViewPoint Latent Action Moel (MVP-LAM), which learns latent actions that are highly informative about ground-truth actions from multi-view videos. MVP-LAM trains latent actions with a cross-viewpoint reconstruction objective, so that a latent action from one view must explain the future in another view, reducing reliance on viewpoint-specific cues. On Bridge V2, MVP-LAM produces more action-centric latent actions, achieving higher mutual information with ground-truth actions and improved action prediction, including under out-of-distribution evaluation. Finally, pretraining VLAs with MVP-LAM latent actions improves downstream manipulation performance on various benchmarks. The code and trained checkpoints are available at https://jmsnu.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。