用3D视觉先验提升机器人视觉控制的泛化能力
VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- 结合3D感知模型与本体感觉,增强视觉空间理解
- 在MetaWorld上优于DP和DP3,长任务表现更优
- 适合需要高精度和强泛化的机器人控制场景
视觉模仿学习框架使机器人能从专家示范中学习操作技能。现有方法多关注策略设计,却常忽略视觉编码器的结构与容量,限制了空间理解与泛化能力。受生物视觉系统启发,该工作提出VGGT-DP,一种融合预训练3D感知模型几何先验与本体感觉反馈的视听觉策略框架。采用视觉几何奠基变压器(VGGT)作为视觉编码器,并引入本体感觉引导的视觉学习策略,对齐感知与内部机器人状态,提升空间定位与闭环控制性能。为降低推理延迟,设计帧级标记复用(FTR)机制,复用重叠观测帧的缓存聚合标记,仅计算最新帧特征,显著减少冗余视觉编码。进一步引入随机标记剪枝以增强策略鲁棒性、降低过拟合。在挑战性的MetaWorld任务上,VGGT-DP显著优于强基线如DP和DP3,尤其在高精度要求和长时程场景中表现突出。
原文摘要 · Abstract (English)
Visual imitation learning frameworks allow robots to learn manipulation skills from expert demonstrations. While existing approaches mainly focus on policy design, they often neglect the structure and capacity of visual encoders, limiting spatial understanding and generalization. Inspired by biological vision systems, which rely on both visual and proprioceptive cues for robust control, we propose VGGT-DP, a visuomotor policy framework that integrates geometric priors from a pretrained 3D perception model with proprioceptive feedback. We adopt the Visual Geometry Grounded Transformer (VGGT) as the visual encoder and introduce a proprioception-guided visual learning strategy to align perception with internal robot states, improving spatial grounding and closed-loop control. To reduce inference latency, we design a frame-wise token reuse (FTR) mechanism that reuses cached VGGT aggregator tokens from overlapping observation frames and computes features only for the latest frame, substantially reducing redundant visual encoding. We further apply random token pruning to enhance policy robustness and reduce overfitting. Experiments on challenging MetaWorld tasks show that VGGT-DP significantly outperforms strong baselines such as DP and DP3, particularly in precision-critical and long-horizon scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。