arXiv:2410.09072cs.RO2024-10被引 1

人类在真实场景中实时指导机器人修复感知失败,高效提升其抓取能力。

iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception

  • 人类通过短时互动暴露物体关键姿态,生成带标注的视觉数据
  • 仅需标注交互末帧,利用半监督策略实现全序列密集标注
  • 少量失败样本即可显著提升模型在复杂场景的分割与抓取成功率

机器人感知模型在真实环境部署时常因分布外条件(如遮挡、杂乱、新物体)而失效。现有方法依赖离线数据收集与重训练,效率低且无法处理运行时故障。本文提出iTeach,一种基于故障驱动的共位交互教学框架,可在真实场景中实时适应机器人感知。协同的人类观察模型预测,识别失败案例,并通过短时人-物交互(HumanPlay)暴露具有信息量的物体配置,同时记录RGB-D序列。为减少标注负担,iTeach采用少样本半监督(FS3)标注策略:仅需用无手眼动与语音命令标注交互末帧,再将标签传播至整段视频,实现密集监督。收集的故障驱动样本用于迭代微调,实现感知模型的渐进式部署时自适应。我们在未见物体实例分割(UOIS)任务上评估,从预训练MSMFormer模型出发,仅用少量故障样本即显著提升不同真实场景下的分割性能。该改进直接带来SceneReplica基准和真实机器人实验中抓取与拾放成功率的提高。结果表明,故障驱动的共位交互教学能高效实现机器人感知的野外自适应,并提升下游操作性能。

原文摘要 · Abstract (English)

Robotic perception models often fail when deployed in real-world environments due to out-of-distribution conditions such as clutter, occlusion, and novel object instances. Existing approaches address this gap through offline data collection and retraining, which are slow and do not resolve deployment-time failures. We propose iTeach, a failure-driven interactive teaching framework for adapting robot perception in the wild. A co-located human observes model predictions during deployment, identifies failure cases, and performs short human-object interaction (HumanPlay) to expose informative object configurations while recording RGB-D sequences. To minimize annotation effort, iTeach employs a Few-Shot Semi- Supervised (FS3) labeling strategy, where only the final frame of a short interaction sequence is annotated using hands-free eye-gaze and voice commands, and labels are propagated across the video to produce dense supervision. The collected failure-driven samples are used for iterative fine-tuning, enabling progressive deployment-time adaptation of the perception model. We evaluate iTeach on unseen object instance segmentation (UOIS) starting from a pretrained MSMFormer model. Using a small number of failure-driven samples, our method significantly improves segmentation performance across diverse real-world scenes. These improvements directly translate to higher grasping and pick-and-place success on the SceneReplica benchmark and real robotic experiments. Our results demonstrate that failure-driven, co-located interactive teaching enables efficient in-the-wild adaptation of robot perception and improves downstream manipulation performance. Project page at https://irvlutd.github.io/iTeach

机器人感知交互教学故障驱动实时适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。