arXiv:2608.04633cs.RO2026-08

让机器人理解指令中的目标物体3D结构,提升精准操作能力

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

论文配图:Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 根据语言指令定位目标物,对齐其专属3D特征
  • 在真实机器人任务中达54%成功率,比基线高26个百分点
  • 适合需要精细抓取和遮挡场景的具身智能应用

近期视觉-语言-动作(VLA)模型通过与三维场景几何对齐提升泛化能力,但这些方法本质上忽视指令:它们统一对齐整个场景,忽略了语言指令指定的目标物体的三维结构。这导致在细粒度操作和目标遮挡任务中失败,因为成功依赖于对目标物体的精确3D理解而非整个场景。为解决此问题,我们提出Mind-VLA,一种面向指令的时空表征对齐方法。首先识别语言指令中的目标物体,生成其规范视图,并提取对应的VAE与VGGT特征;随后将VLA模型的潜在表示与这些特征对齐,实现指令感知的3D理解。Mind-VLA在LIBERO数据集上达到94.4%成功率,在CALVIN上达到4.47分,仅使用345M参数的紧凑主干网络。在包含目标遮挡的真实机器人任务中,平均成功率达54%,优于匹配的场景-VGGT对照组26个百分点。

原文摘要 · Abstract (English)

Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-agnostic: the representations align the entire scene uniformly, neglecting the 3D geometry of the specific target object designated by the language instruction. This causes failures on fine-grained manipulation and target occlusion tasks, where success depends on accurate 3D understanding of the target object rather than the entire scene. To address this, we present Mind-VLA, an instruction-aware spatial representation alignment method for VLA models. Specifically, Mind-VLA first obtains the target object specified by the language instruction, then prepares its canonical target views and extracts the corresponding VAE and VGGT features. Finally, the latent representation of the VLA model is aligned with these features to enable instruction-aware 3D understanding. Mind-VLA reaches 94.4% on LIBERO and 4.47 on CALVIN with a compact 345M-parameter backbone. On real-robot tasks with target occlusion, Mind-VLA reaches 54% average success, outperforming the matched scene-VGGT control by 26 percentage points.

视觉语言动作目标定位具身智能3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。