arXiv:2602.23583cs.RO2026-02被引 1

用点击代替语言指令,精准操控复杂环境中的物体。

VCA: Vision-Click-Action Framework for Precise Manipulation of Segmented Objects in Target Ambiguous Environments

  • 通过点击图像直接指定目标物体,替代繁琐语言描述。
  • 在多相似物体环境中实现高精度实例级操作,准确率显著提升。
  • 适合需要快速、低负担人机交互的工业机器人场景。

视觉-点击-动作(VCA)框架旨在解决视觉-语言-动作(VLA)模型依赖语言带来的模糊性、认知负担及精确目标识别难题,尤其在存在多个视觉相似物体的环境中。该框架引入预训练分割模型,以直接的点击式视觉交互取代冗长的文本指令。操作者可在机器人的二维相机视图中通过点击明确指定目标物体,有效减少理解误差,降低认知负荷,并提供一种实用且可扩展的现实机器人操作接口。实验验证了VCA在指定目标物体上的实例级操纵有效性。演示视频见https://robrosinc.github.io/vca/。

原文摘要 · Abstract (English)

The reliance on language in Vision-Language-Action (VLA) models introduces ambiguity, cognitive overhead, and difficulties in precise object identification and sequential task execution, particularly in environments with multiple visually similar objects. To address these limitations, we propose Vision-Click-Action (VCA), a framework that replaces verbose textual commands with direct, click-based visual interaction using pretrained segmentation models. By allowing operators to specify target objects clearly through visual selection in the robot's 2D camera view, VCA reduces interpretation errors, lowers cognitive load, and provides a practical and scalable alternative to language-driven interfaces for real-world robotic manipulation. Experimental results validate that the proposed VCA framework achieves effective instance-level manipulation of specified target objects. Experiment videos are available at https://robrosinc.github.io/vca/.

机器人操控视觉交互点击输入实例分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。