arXiv:2411.11409cs.CVcs.AI2024-11NeurIPS被引 15

将家具组装视频与说明书对齐,实现时空精准理解。

IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos

  • 构建跨模态数据集,对齐3D零件、说明书与互联网视频
  • 在视频中定位零件动作,支持组装计划生成等5项任务
  • 解决遮挡、视角变化等挑战,适合视觉理解研究者

家具组装是日常生活中常见的任务,用于构建如宜家家具等复杂三维结构。尽管自主装配智能体已有显著进展,但现有数据集尚未解决视频中组装指令的4D(空间+时间)对齐问题,这限制了对组装过程的全面理解。我们提出IKEA Video Manuals数据集,包含家具零件的3D模型、说明书、来自互联网的组装视频,以及这些多模态数据之间的密集时空对齐标注。为验证该数据集的实用性,我们展示了五项关键应用:组装计划生成、基于零件的分割、基于零件的姿态估计、视频目标分割,以及基于说明书视频的家具组装。每项任务均提供评估指标和基线方法。在标注数据上的实验揭示了在视频中对齐组装指令的诸多挑战,包括遮挡、视角变化和长序列组装过程。

原文摘要 · Abstract (English)

Shape assembly is a ubiquitous task in daily life, integral for constructing complex 3D structures like IKEA furniture. While significant progress has been made in developing autonomous agents for shape assembly, existing datasets have not yet tackled the 4D grounding of assembly instructions in videos, essential for a holistic understanding of assembly in 3D space over time. We introduce IKEA Video Manuals, a dataset that features 3D models of furniture parts, instructional manuals, assembly videos from the Internet, and most importantly, annotations of dense spatio-temporal alignments between these data modalities. To demonstrate the utility of IKEA Video Manuals, we present five applications essential for shape assembly: assembly plan generation, part-conditioned segmentation, part-conditioned pose estimation, video object segmentation, and furniture assembly based on instructional video manuals. For each application, we provide evaluation metrics and baseline methods. Through experiments on our annotated data, we highlight many challenges in grounding assembly instructions in videos to improve shape assembly, including handling occlusions, varying viewpoints, and extended assembly sequences.

4D对齐视频理解家具组装多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。