让机器人主动调整视角和焦距,实现高精度3D抓取
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
- 分两阶段:先定位关键区域,再动态选最优视角并放大细节
- 在3个仿真环境上超越当前最佳基线,真实场景也有效
- 适合需要精细操作的机器人任务,如拆解或装配
近期机器人操作研究利用预训练视觉语言模型,并引入3D空间信号以提升动作预测能力,形成有前景的视觉-语言-动作(VLA)范式。然而,现有方法普遍忽视主动感知的重要性:通常依赖固定安装于手腕的摄像头,提供以末端执行器为中心的视角。这导致模型无法在任务执行中自适应选择最优视角或分辨率,显著限制了其在长周期任务和精细操作场景中的表现。为此,我们提出ActiveVLA,一种赋予机器人主动感知能力的新颖视觉-语言-动作框架,用于高精度、细粒度的3D操作。ActiveVLA采用粗到精策略,分为两阶段:(1) 关键区域定位。将3D输入投影至多视角2D图像,识别关键3D区域,支持动态空间感知;(2) 主动感知优化。基于定位结果,采用主动视角选择策略,选取最大化模态相关性与多样性、同时最小化遮挡的最优视角,并对关键区域进行3D缩放以提升分辨率。上述步骤共同实现更精细的主动感知,支持精准操作。大量实验表明,ActiveVLA在三个仿真基准上实现精确3D操作,性能优于当前最先进基线。此外,该方法可无缝迁移至真实世界,使机器人在复杂环境中学习高精度任务。
原文摘要 · Abstract (English)
Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising vision-language-action (VLA) paradigm. However, most existing approaches overlook the importance of active perception: they typically rely on static, wrist-mounted cameras that provide an end-effector-centric viewpoint. As a result, these models are unable to adaptively select optimal viewpoints or resolutions during task execution, which significantly limits their performance in long-horizon tasks and fine-grained manipulation scenarios. To address these limitations, we propose ActiveVLA, a novel vision-language-action framework that empowers robots with active perception capabilities for high-precision, fine-grained manipulation. ActiveVLA adopts a coarse-to-fine paradigm, dividing the process into two stages: (1) Critical region localization. ActiveVLA projects 3D inputs onto multi-view 2D projections, identifies critical 3D regions, and supports dynamic spatial awareness. (2) Active perception optimization. Drawing on the localized critical regions, ActiveVLA uses an active view selection strategy to choose optimal viewpoints. These viewpoints aim to maximize amodal relevance and diversity while minimizing occlusions. Additionally, ActiveVLA applies a 3D zoom-in to improve resolution in key areas. Together, these steps enable finer-grained active perception for precise manipulation. Extensive experiments demonstrate that ActiveVLA achieves precise 3D manipulation and outperforms state-of-the-art baselines on three simulation benchmarks. Moreover, ActiveVLA transfers seamlessly to real-world scenarios, enabling robots to learn high-precision tasks in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。