让视觉语言动作模型学会空间几何理解,提升操作任务成功率。
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- 用几何知识蒸馏将3D空间信息融入多模态表征
- 在LIBERO上达94.1%成功率,优于无几何先验的基线
- 无需深度传感器或3D标注,适合真实机器人应用
现有视觉-语言-动作(VLA)模型主要依赖RGB信息,忽略对空间推理和操作至关重要的几何线索。本文提出GLaD,一种在预训练阶段通过知识蒸馏引入3D几何先验的几何感知VLA框架。不同于仅将几何特征蒸馏至视觉编码器,我们对齐大语言模型中对应视觉标记的隐藏状态与冻结的几何感知视觉变压器(VGGT)特征,确保几何理解深度融入驱动动作预测的多模态表征。在包含4个LIBERO任务套件的Bridge数据集上预训练后,GLaD在所有任务上平均成功率达94.1%,优于使用相同预训练数据的UniVLA(92.5%)。结果验证了几何感知预训练可增强空间推理能力与策略泛化性,且无需显式深度传感器或3D标注。
原文摘要 · Abstract (English)
Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA framework that incorporates 3D geometric priors during pretraining through knowledge distillation. Rather than distilling geometric features solely into the vision encoder, we align the LLM's hidden states corresponding to visual tokens with features from a frozen geometry-aware vision transformer (VGGT), ensuring that geometric understanding is deeply integrated into the multimodal representations that drive action prediction. Pretrained on the Bridge dataset with this geometry distillation mechanism, GLaD achieves 94.1% average success rate across four LIBERO task suites, outperforming UniVLA (92.5%) which uses identical pretraining data. These results validate that geometry-aware pretraining enhances spatial reasoning and policy generalization without requiring explicit depth sensors or 3D annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。