arXiv:2602.20231cs.ROcs.CV2026-02被引 3

用深度信息增强视觉动作表征,提升机器人操作精度

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

  • 融合RGB与深度图的统一潜空间学习,显式建模跨模态交互
  • 在仿真与真实场景中,新方法在多种任务上均超越基线模型
  • 适合需要精准空间理解的机器人操控研究者

从无标签视频中学习的潜动作表征已成为无需机器人动作监督的视觉-语言-动作(VLA)模型预训练的有力范式。然而,仅基于RGB观测得到的潜动作主要编码外观驱动的动力学,缺乏对精确操作至关重要的3D几何结构。为此,我们提出UniLACT,一种基于Transformer的VLA模型,通过深度感知的潜预训练引入几何结构,使下游策略继承更强的空间先验。为此,我们设计了UniLARN,一个基于逆向与前向动力学目标的统一潜动作学习框架,联合学习RGB与深度的共享嵌入空间,并显式建模其跨模态交互。该框架生成模态特定及统一的潜动作表示,作为UniLACT深度感知预训练的伪标签。大量仿真与真实世界实验表明,深度感知的统一潜动作表示有效。UniLACT在同域与跨域预训练、已见与未见操作任务中持续优于基于RGB的潜动作基线。

原文摘要 · Abstract (English)

Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-action (VLA) models without explicit robot action supervision. However, latent actions derived solely from RGB observations primarily encode appearance-driven dynamics and lack explicit 3D geometric structure, which is essential for precise and contact-rich manipulation. To address this limitation, we introduce UniLACT, a transformer-based VLA model that incorporates geometric structure through depth-aware latent pretraining, enabling downstream policies to inherit stronger spatial priors. To facilitate this process, we propose UniLARN, a unified latent action learning framework based on inverse and forward dynamics objectives that learns a shared embedding space for RGB and depth while explicitly modeling their cross-modal interactions. This formulation produces modality-specific and unified latent action representations that serve as pseudo-labels for the depth-aware pretraining of UniLACT. Extensive experiments in both simulation and real-world settings demonstrate the effectiveness of depth-aware unified latent action representations. UniLACT consistently outperforms RGB-based latent action baselines under in-domain and out-of-domain pretraining regimes, as well as on both seen and unseen manipulation tasks.The project page is at https://manishgovind.github.io/unilact-vla/

视觉语言动作潜动作深度感知机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。