通过对比提示调度,让机器人在不同传感器配置下自适应地完成视觉动作任务。
Learning Adaptive Cross-Embodiment Visuomotor Policy with Contrastive Prompt Orchestration
- 用视觉、动作和文本三重对比学习构建可调提示池,捕捉细微环境差异。
- 动态组合提示以实时生成最优状态表征,提升样本效率和泛化能力。
- 零样本迁移效果好,适合跨设备、跨场景的机器人控制应用。
为具身智能体学习自适应视觉运动策略仍面临巨大挑战,尤其在传感器配置和动态属性差异较大的跨具身场景中。传统方法难以区分任务相关特征与领域特异性变化(如光照、视角、旋转),导致样本效率低且在未见环境中出现灾难性失败。为此,我们提出对比提示调度(CAPO)方法,融合对比提示学习与自适应提示编排机制。通过结合视觉、时间动作与文本目标的混合对比学习,构建可学习提示池,每个提示编码精细的领域因子。基于这些提示,引入自适应提示编排机制,根据当前观测动态聚合提示,使智能体能即时识别主导领域因素并构建最优状态表示。这有效屏蔽了无关干扰,避免对源域过拟合。大量实验表明,CAPO在样本效率和最终性能上显著优于现有基线;关键是在光照、视角、旋转等剧烈变化的未见目标领域中展现出优异的零样本适应能力,验证了其作为跨具身视觉运动策略适应的有效解决方案。
原文摘要 · Abstract (English)
Learning adaptive visuomotor policies for embodied agents remains a formidable challenge, particularly when facing cross-embodiment variations such as diverse sensor configurations and dynamic properties. Conventional learning approaches often struggle to separate task-relevant features from domain-specific variations (e.g., lighting, field-of-view, and rotation), leading to poor sample efficiency and catastrophic failure in unseen environments. To bridge this gap, we propose ContrAstive Prompt Orchestration (CAPO), a novel approach for learning visuomotor policies that integrates contrastive prompt learning and adaptive prompt orchestration. For prompt learning, we devise a hybrid contrastive learning strategy that integrates visual, temporal action, and text objectives to establish a pool of learnable prompts, where each prompt induces a visual representation encapsulating fine-grained domain factors. Based on these learned prompts, we introduce an adaptive prompt orchestration mechanism that dynamically aggregates these prompts conditioned on current observations. This enables the agent to adaptively construct optimal state representations by identifying dominant domain factors instantaneously. Consequently, the policy optimization is effectively shielded from irrelevant interference, preventing the common issue of overfitting to source domains. Extensive experiments demonstrate that CAPO significantly outperforms state-of-the-art baselines in sample efficiency and asymptotic performance. Crucially, it exhibits superior zero-shot adaptation across unseen target domains characterized by drastic environmental (e.g., illumination) and physical shifts (e.g., field-of-view and rotation), validating its effectiveness as a viable solution for cross-embodiment visuomotor policy adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。