arXiv:2502.03270cs.ROcs.AI2025-02被引 3

预训练视觉模型在时序任务中存在时间混淆问题,影响机器人决策效果。

The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning

  • 提出通过解耦时间信息来缓解视觉表征中的时间纠缠
  • 发现策略成功率与潜在空间捕捉任务进展线索的能力强相关
  • 适合做视觉-运动策略学习的算法研究者参考

预训练视觉表示(PVRs)在视觉-运动策略学习中已取得显著进展,但如何有效利用仍具挑战。本文识别出一个关键内在问题:在序列决策任务中使用时间不变的预训练模型会引发时间纠缠。这是因为PVRs针对静态图像理解进行优化,难以表征视觉运动控制所需的时间依赖性。本文量化了时间纠缠的影响,证明策略成功率与潜在空间捕捉任务进展线索的能力呈强相关。基于此,提出一种简单但有效的解耦基线方法以缓解时间纠缠。实验表明,仅通过增加时间特征的方法不足以解决问题,凸显了显式处理时间解耦对鲁棒视觉-运动策略学习的必要性。

原文摘要 · Abstract (English)

The integration of pre-trained visual representations (PVRs) has significantly advanced visuomotor policy learning. However, effectively leveraging these models remains a challenge. We identify temporal entanglement as a critical, inherent issue when using these time-invariant models in sequential decision-making tasks. This entanglement arises because PVRs, optimised for static image understanding, struggle to represent the temporal dependencies crucial for visuomotor control. In this work, we quantify the impact of temporal entanglement, demonstrating a strong correlation between a policy's success rate and the ability of its latent space to capture task-progression cues. Based on these insights, we propose a simple, yet effective disentanglement baseline designed to mitigate temporal entanglement. Our empirical results show that traditional methods aimed at enriching features with temporal components are insufficient on their own, highlighting the necessity of explicitly addressing temporal disentanglement for robust visuomotor policy learning.

视觉-运动时间解耦预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。