从第一视角视频中根据指令生成6自由度物体操作轨迹
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
- 利用全局构建的多视角视频数据集提取大规模操作轨迹
- 在HOT3D数据集上成功生成有效6DoF轨迹,验证方法可行性
- 适合研究具身智能与视觉语言模型交互的学者参考
学习在常见场景中使用工具或物体,特别是根据指令以多种方式操作它们,是发展交互式机器人的关键挑战。训练模型生成此类操作轨迹需要大量且多样化的详细操作示范,这在大规模上几乎无法实现。本文提出一个框架,利用全球努力构建的大规模内外视角视频数据集Exo-Ego4D,大规模提取多样化操作轨迹。基于这些带有文本动作描述的轨迹,我们开发了基于视觉和点云的言语模型来生成轨迹。在最近提出的基于第一人称视觉的高质量轨迹数据集HOT3D上,我们确认所提模型能够成功生成有效物体轨迹,为从第一人称视觉中根据动作描述生成6自由度操作轨迹这一新任务建立了训练数据集和基线模型。
原文摘要 · Abstract (English)
Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large and diverse collection of detailed manipulation demonstrations for various objects, which is nearly unfeasible to gather at scale. In this paper, we propose a framework that leverages large-scale ego- and exo-centric video datasets -- constructed globally with substantial effort -- of Exo-Ego4D to extract diverse manipulation trajectories at scale. From these extracted trajectories with the associated textual action description, we develop trajectory generation models based on visual and point cloud-based language models. In the recently proposed egocentric vision-based in-a-quality trajectory dataset of HOT3D, we confirmed that our models successfully generate valid object trajectories, establishing a training dataset and baseline models for the novel task of generating 6DoF manipulation trajectories from action descriptions in egocentric vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。