arXiv:2512.20014cs.ROcs.AI2025-12被引 6

让机器人通过几张照片认出用户特定物品并执行指令。

Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting

  • 用参考图构建视觉记忆,通过检测与匹配定位目标物
  • 在不训练模型的前提下提升识别准确率,成功率达92.3%
  • 适合需要个性化交互的家用机器人场景

尽管视觉-语言-动作(VLA)模型对通用指令泛化良好,但在处理如'拿我的杯子'这类个性化指令时表现不佳,因需从外观相似物体中精准识别特定实例。本文研究个人物品操作任务,即在仅提供少量参考图像的情况下,让冻结的VLA模型识别并操控用户指定物体。为此提出视觉注意力提示(VAP),一种无需训练的感知适配器:将参考图像视为非参数化视觉记忆,利用开放词汇检测和基于嵌入的匹配在场景中定位目标物,并通过突出显示该物和重写指令的方式注入视觉提示。构建了两个仿真基准(Personalized-SIMPLER 和 Personalized-VLABench)及一个真实桌面环境基准,评估多机器人、多任务下的个性化操作性能。实验表明,VAP在成功率和正确操作率上均显著优于通用策略和基于标记学习的基线方法,有效弥合语义理解与实例级控制之间的差距。

原文摘要 · Abstract (English)

While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one specific instance among visually similar objects. We study this setting of manipulating personal objects, in which a VLA must identify and control a user-specific object unseen during training using only a few reference images. To address this challenge, we propose Visual Attentive Prompting (VAP), a simple-yet-effective training-free perceptual adapter that equips frozen VLAs with top-down selective attention. VAP treats the reference images as a non-parametric visual memory, grounds the personal object in the scene through open-vocabulary detection and embedding-based matching, and then injects this grounding as a visual prompt by highlighting the object and rewriting the instruction. We construct two simulation benchmarks, Personalized-SIMPLER and Personalized-VLABench, and a real-world tabletop benchmark to evaluate personalized manipulation across multiple robots and tasks. Experiments show that VAP consistently outperforms generic policies and token-learning baselines in both success rate and correct-object manipulation, helping to bridge the gap between semantic understanding and instance-level control.

机器人操作个性化视觉提示VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。