让机器人通过跨模态预测提前看清动作后果,提升操作成功率
CLaD: Planning with Grounded Foresight via Cross-Modal Latent Dynamics
- 用不对称交叉注意力建模动作下本体与语义状态的联合演化
- 在LIBERO-LONG上达94.7%成功率,参数量远少于大模型
- 适合追求高效精准决策的机器人系统研发者
机器人操作涉及本体和语义状态的内在耦合变化。现有方法通常在语义或潜在空间中规划,未显式对齐跨模态转变。为此,我们提出CLaD框架,通过非对称交叉注意力建模动作下本体与语义状态的联合演化,使运动变化可查询语义信息。CLaD利用自监督目标与EMA目标编码器预测具身化潜在前瞻,并辅以辅助重建损失,防止表征坍塌,同时将预测锚定在可观测状态。预测结果结合实时观测,调节扩散策略生成动作。在LIBERO-LONG基准上,CLaD实现94.7%的成功率,性能媲美大型视觉-语言模型,但参数量显著更少。
原文摘要 · Abstract (English)
Robotic manipulation involves kinematic and semantic transitions that are inherently coupled via underlying actions. However, existing approaches plan within either semantic or latent space without explicitly aligning these cross-modal transitions. To address this, we propose CLaD, a framework that models how proprioceptive and semantic states jointly evolve under actions through asymmetric cross-attention that allows kinematic transitions to query semantic ones. CLaD predicts grounded latent foresights via self-supervised objectives with EMA target encoders and auxiliary reconstruction losses, preventing representation collapse while anchoring predictions to observable states. Predicted foresights are modulated with observations to condition a diffusion policy for action generation. On LIBERO-LONG benchmark, CLaD achieves 94.7\% success rate, competitive with large VLAs with significantly fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。