arXiv:2603.12751cs.CVcs.LG2026-03

用人类示范视频自动训练机器人识别新物体,无需语言描述。

Show, Don't Tell: Detecting Novel Objects by Watching Human Videos

  • 用示范视频自动生成训练数据,直接让模型看物体
  • 在真实机器人上实现新物体检测,准确率显著提升
  • 适合需要快速适应新物品的机器人场景

机器人如何在人类示范过程中快速识别并认出新出现的物体?现有闭集目标检测器常因物体分布外而失效。尽管开放集检测器(如视觉语言模型)有时有效,但通常需昂贵且繁琐的人工提示工程来唯一识别新物体实例。本文提出一种自监督系统,通过人类示范视频本身自动构建训练数据,无需复杂的语言描述和人工提示工程,即可训练专属的目标检测器。在“看,不要说”(Show, Don't Tell)范式中,我们直接向检测器展示目标物体,而非用复杂语言描述它们。该方法绕过语言环节,可快速训练针对特定示范中物体的专用检测器。我们开发了一个集成于机器人的系统,实现在真实机器人上部署自动数据生成与新物体检测。实验结果表明,该流程在操作物体的检测与识别上显著优于现有最先进方法,从而提升了机器人的任务完成能力。

原文摘要 · Abstract (English)

How can a robot quickly identify and recognize new objects shown to it during a human demonstration? Existing closed-set object detectors frequently fail at this because the objects are out-of-distribution. While open-set detectors (e.g., VLMs) sometimes succeed, they often require expensive and tedious human-in-the-loop prompt engineering to uniquely recognize novel object instances. In this paper, we present a self-supervised system that eliminates the need for tedious language descriptions and expensive prompt engineering by training a bespoke object detector on an automatically created dataset, supervised by the human demonstration itself. In our approach, "Show, Don't Tell," we show the detector the specific objects of interest during the demonstration, rather than telling the detector about these objects via complex language descriptions. By bypassing language altogether, this paradigm enables us to quickly train bespoke detectors tailored to the relevant objects observed in human task demonstrations. We develop an integrated on-robot system to deploy our "Show, Don't Tell" paradigm of automatic dataset creation and novel object-detection on a real-world robot. Empirical results demonstrate that our pipeline significantly outperforms state-of-the-art detection and recognition methods for manipulated objects, leading to improved task completion for the robot.

机器人物体识别自监督示范学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。