arXiv:2509.26004cs.CVcs.AI2025-09

用人类口述描述弱监督训练手部操作物体分割模型。

Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations

  • 通过人类口述动作描述学习手部物体关联,无需精细标注。
  • 在EPIC-Kitchens和Ego4D上达到全监督方法50%以上性能。
  • 适合无标注数据的视觉交互研究者使用。

从第一人称视角图像中进行像素级物体识别,可支持辅助技术、工业安全和行为监控等关键应用。然而,当前进展受限于标注数据稀缺,现有方法依赖昂贵的人工标注。本文提出利用口述描述——佩戴相机者自然叙述的动作内容,其中包含被操作物体的线索——来学习人机交互检测。我们引入了口述监督的手部物体分割(NS-iHOS)新任务,模型需在弱监督下从自然语言口述中学习分割手部物体,且推理时不再使用口述。我们提出端到端的WISH模型,通过蒸馏口述信息学习合理的手物关联,实现无需口述即可完成手部物体分割。在EPIC-Kitchens和Ego4D上的实验表明,WISH超越所有基线,性能超过全监督方法的50%,且无需细粒度像素标注。代码与数据见https://fpv-iplab.github.io/WISH。

原文摘要 · Abstract (English)

Pixel-level recognition of objects manipulated by the user from egocentric images enables key applications spanning assistive technologies, industrial safety, and activity monitoring. However, progress in this area is currently hindered by the scarcity of annotated datasets, as existing approaches rely on costly manual labels. In this paper, we propose to learn human-object interaction detection leveraging narrations $\unicode{x2013}$ natural language descriptions of the actions performed by the camera wearer which contain clues about manipulated objects. We introduce Narration-Supervised in-Hand Object Segmentation (NS-iHOS), a novel task where models have to learn to segment in-hand objects by learning from natural-language narrations in a weakly-supervised regime. Narrations are then not employed at inference time. We showcase the potential of the task by proposing Weakly-Supervised In-hand Object Segmentation from Human Narrations (WISH), an end-to-end model distilling knowledge from narrations to learn plausible hand-object associations and enable in-hand object segmentation without using narrations at test time. We benchmark WISH against different baselines based on open-vocabulary object detectors and vision-language models. Experiments on EPIC-Kitchens and Ego4D show that WISH surpasses all baselines, recovering more than 50% of the performance of fully supervised methods, without employing fine-grained pixel-wise annotations. Code and data can be found at https://fpv-iplab.github.io/WISH.

弱监督第一人称视觉物体分割自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。