用视觉标注提升语言指令的物体指代精度,让机器人更懂人话。
Point What You Mean: Visually Grounded Instruction Policy
- 在语言指令中加入框选等视觉线索,精准定位目标物体。
- 在杂乱或陌生场景中,性能显著优于纯文本指令模型。
- 自动标注数据流水线大幅降低人工成本,适合实际部署。
视觉-语言-动作(VLA)模型将视觉与语言结合用于具身控制,但在仅依赖文本指令时,物体指代能力受限,尤其在杂乱或分布外(OOD)场景中。本研究提出Point-VLA,一种可即插即用的策略,通过添加显式视觉线索(如边界框)来消除指代歧义,实现像素级物体定位。为高效扩展视觉标注数据集,我们还开发了低人工成本的自动化数据标注流程。在多种真实世界指代任务上评估表明,Point-VLA在杂乱或未见物体场景中均显著优于纯文本指令的VLA模型,表现出强泛化能力。结果证明,通过像素级视觉定位,Point-VLA能有效解决物体指代歧义,实现更通用的具身控制。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (OOD) scenes. In this study, we introduce the Point-VLA, a plug-and-play policy that augments language instructions with explicit visual cues (e.g., bounding boxes) to resolve referential ambiguity and enable precise object-level grounding. To efficiently scale visually grounded datasets, we further develop an automatic data annotation pipeline requiring minimal human effort. We evaluate Point-VLA on diverse real-world referring tasks and observe consistently stronger performance than text-only instruction VLAs, particularly in cluttered or unseen-object scenarios, with robust generalization. These results demonstrate that Point-VLA effectively resolves object referring ambiguity through pixel-level visual grounding, achieving more generalizable embodied control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。