提出AcTOL方法,让视觉语言模型学会动作的自然时序与连续性。
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
- 用帧间语义差异对比实现动作时序建模
- 引入布朗桥约束保证中间帧平滑过渡
- 提升机器人对不同指令风格的泛化能力
在人类动作视频上预训练视觉-语言表示,已成为减少对大规模专家示范依赖的有前景方法。然而,以往方法多采用基于目标达成启发式的时序对比学习,逐步对齐从初始帧到最终帧的语言指令,这种对后续帧的过度关注可能导致错误的视觉-语言关联,因为动作可能提前结束或包含无关结尾内容。为此,我们提出动作时序一致性学习(AcTOL),在无严格目标约束下学习有序且连续的视觉-语言表示。AcTOL将视频视为连续轨迹,通过(1)对比帧间语义差异以反映自然顺序,(2)施加局部布朗桥约束确保中间帧间的平滑过渡。在仿真和真实机器人上的广泛模仿学习实验表明,预训练特征显著提升了下游操作任务性能,且对不同语言风格的指令具有高度鲁棒性,为通用具身智能体提供可行路径。
原文摘要 · Abstract (English)
Pre-training vision-language representations on human action videos has emerged as a promising approach to reduce reliance on large-scale expert demonstrations for training embodied agents. However, prior methods often employ time contrastive learning based on goal-reaching heuristics, progressively aligning language instructions from the initial to the final frame. This overemphasis on future frames can result in erroneous vision-language associations, as actions may terminate early or include irrelevant moments in the end. To address this issue, we propose Action Temporal Coherence Learning (AcTOL) to learn ordered and continuous vision-language representations without rigid goal-based constraint. AcTOL treats a video as a continuous trajectory where it (1) contrasts semantic differences between frames to reflect their natural ordering, and (2) imposes a local Brownian bridge constraint to ensure smooth transitions across intermediate frames. Extensive imitation learning experiments on both simulated and real robots show that the pretrained features significantly enhance downstream manipulation tasks with high robustness to different linguistic styles of instructions, offering a viable pathway toward generalized embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。