通过融合几何先验提升视觉语言动作模型的视角泛化能力。
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- 用冻结的几何视觉模型提取特征,避免从头学习3D一致性。
- 在LIBERO和CALVIN上实现平均35%和超11%的未见视角成功率提升。
- 适用于真实机器人平台,对连续与离散动作空间均有效。
视觉-语言-动作(VLA)模型在未见摄像头视角下常表现不佳,根源在于难以从2D图像中推断出可靠的3D几何结构。本文提出GeoAware-VLA,通过将强几何先验融入视觉主干网络,提升视角不变性。不训练视觉编码器也不依赖显式3D数据,而是使用一个冻结的预训练几何视觉模型作为特征提取器,再通过轻量级可训练投影层将这些富含几何信息的特征适配至策略解码器,从而减轻其学习3D一致性的负担。在LIBERO和CALVIN基准上的大量实验表明,GeoAware-VLA在保持甚至提升分布内性能的同时,在零样本泛化到未见相机姿态方面取得显著进步:相比基线模型,LIBERO上未见视角成功率平均提升35个百分点,CALVIN上提升超过11个百分点。关键的是,这些改进可迁移到真实机器人平台,表现出显著性能提升。该方法在连续与离散动作空间中均有效,表明鲁棒的几何定位是构建更通用机器人智能体的关键要素。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models often fail to generalize to unseen camera viewpoints, a limitation stemming from their difficulty in inferring robust 3D geometry from 2D images. We introduce GeoAware-VLA, a simple yet effective approach that enhances viewpoint invariance by integrating strong geometric priors into the vision backbone. Instead of training a visual encoder or relying on explicit 3D data, we leverage a frozen, pretrained geometric vision model as a feature extractor. A lightweight, trainable projection layer then adapts these geometrically-rich features for the policy decoder, relieving it of the burden of learning 3D consistency from scratch. Through extensive evaluations on the LIBERO and CALVIN benchmarks, we show that GeoAware-VLA preserves and even improves in-distribution performance while achieving substantial gains in zero-shot generalization to unseen camera poses, improving unseen-view success rates by an average of 35 percentage points on LIBERO and over 11 percentage points on CALVIN compared to their respective baselines. Crucially, these gains transfer to the physical world, where our model shows significant improvement on a real robotic platform. Our approach proves effective across both continuous and discrete action spaces, highlighting that robust geometric grounding is a key ingredient for building more generalizable robotic agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。