让机器人在杂乱环境中更准地识别目标并执行操作。
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- 将感知与决策分离,用多视角物体中心化和几何信息增强视觉输入。
- 在真实桌面上测试,对干扰物、目标缺失等场景的鲁棒性显著提升。
- 适合需要高精度环境理解的机器人操作任务,如家庭服务或工业分拣。
近期视觉-语言-动作(VLA)模型通过微调大型视觉-语言模型(VLM)实现动作预测,在通用机器人操作上取得进展。然而,多数VLA将感知与控制耦合于单一管道,仅优化动作性能,削弱了语言引导的定位能力。在真实桌面实验中,策略常误抓、受杂乱干扰、过度依赖背景外观。为此,我们提出OBEYED-VLA(OBject-centric and gEometrY groundED VLA),显式解耦感知定位与动作推理。该框架在原始RGB图像上增加感知模块,将多视角输入转化为任务相关的物体中心、几何感知的观测。模块包含基于VLM的物体中心定位阶段,跨摄像头选择任务相关物体区域;以及互补的几何定位阶段,强调物体三维结构而非外观。由此生成的接地视图输入预训练的VLA策略,并仅在无杂乱、单物体演示数据上进行微调。在真实世界UR10e桌面设置中,OBEYED-VLA在四种挑战场景及多个难度级别下显著优于强基线:干扰物、目标缺失拒绝、背景外观变化、未见物体的杂乱操作。消融实验证明,语义定位与几何感知均对性能提升至关重要。结果表明,将感知作为显式的物体中心组件,是增强和泛化基于VLA的机器人操作的有效途径。
原文摘要 · Abstract (English)
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action prediction. Yet most VLAs entangle perception and control in a monolithic pipeline optimized purely for action, which can erode language-conditioned grounding. In our real-world tabletop tests, policies over-grasp when the target is absent, are distracted by clutter, and overfit to background appearance. To address these issues, we propose OBEYED-VLA (OBject-centric and gEometrY groundED VLA), a framework that explicitly disentangles perceptual grounding from action reasoning. Instead of operating directly on raw RGB, OBEYED-VLA augments VLAs with a perception module that grounds multi-view inputs into task-conditioned, object-centric, and geometry-aware observations. This module includes a VLM-based object-centric grounding stage that selects task-relevant object regions across camera views, along with a complementary geometric grounding stage that emphasizes the 3D structure of these objects over their appearance. The resulting grounded views are then fed to a pretrained VLA policy, which we fine-tune exclusively on single-object demonstrations collected without environmental clutter or non-target objects. On a real-world UR10e tabletop setup, OBEYED-VLA substantially improves robustness over strong VLA baselines across four challenging regimes and multiple difficulty levels: distractor objects, absent-target rejection, background appearance changes, and cluttered manipulation of unseen objects. Ablation studies confirm that both semantic grounding and geometry-aware grounding are critical to these gains. Overall, the results indicate that making perception an explicit, object-centric component is an effective way to strengthen and generalize VLA-based robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。