用更高效几何感知视觉编码器提升机器人操作成功率
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
- 提出轻量级几何感知编码器eVGGT,替代传统视觉模型
- 在仿真与真实场景中,操作成功率最高提升6.5%
- 适合追求高精度空间理解的机器人系统部署
现有基于RGB的模仿学习方法通常采用ResNet或ViT等传统视觉编码器,缺乏显式的三维推理能力。近期的几何引导视觉模型(如VGGT)提供了更强的空间理解能力,是解决该问题的有力候选。本文研究将几何感知视觉表示引入机器人操作任务。结果表明,在模拟和真实环境中,将几何感知编码器集成到ACT和DP等模仿学习框架中,相比标准编码器,单手与双手操作任务的成功率最高提升6.5%。尽管如此,多数几何引导模型计算开销大,限制了其在实际机器人系统中的应用。为此,本文提出eVGGT,一种从VGGT蒸馏得到的高效几何感知编码器。eVGGT速度接近VGGT的9倍,体积仅为1/5,同时保留了强大的三维推理能力。代码与预训练模型将公开,以推动几何感知机器人研究。
原文摘要 · Abstract (English)
Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-grounded vision models, such as VGGT~\cite{wang2025vggt}, provide robust spatial understanding and are promising candidates to address this limitation. This work investigates the integration of geometry-aware visual representations into robotic manipulation. Our results suggest that incorporating the geometry-aware vision encoder into imitation learning frameworks, including ACT and DP, yields up to 6.5% improvement over standard vision encoders in success rate across single- and bi-manual manipulation tasks in both simulation and real-world settings. Despite these benefits, most geometry-grounded models require high computational cost, limiting their deployment in practical robotic systems. To address this challenge, we propose eVGGT, an efficient geometry-aware encoder distilled from VGGT. eVGGT is nearly 9 times faster and 5 times smaller than VGGT, while preserving strong 3D reasoning capabilities. Code and pretrained models will be released to facilitate further research in geometry-aware robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。