构建合成场景图数据集,精准标注人与物体交互关系。
HOIverse: A Synthetic Scene Graph Dataset With Human Object Interactions
- 基于参数化关系生成人与物体间的精确交互标注
- 包含RGB图像、分割掩码、深度图和人体关键点等多模态数据
- 适用于人机共存环境下的场景理解研究
当人类与机器人在环境中共存时,场景理解对导航、规划等下游任务至关重要。当前室内场景中缺乏可靠的数据集支持以人类为场景组成部分的视觉理解研究。场景图可提供结构化表示,用于视觉场景分析。为此,我们提出HOIverse——一个融合场景图与人-物体交互的合成数据集,包含人与周围物体间准确且密集的关系标注,以及对应的RGB图像、分割掩码、深度图和人体关键点。通过计算各类物体对及人-物体对之间的参数化关系,实现清晰无歧义的关系定义。同时,在该数据集上对现有最先进的场景图生成模型进行基准测试,以预测参数化关系和人-物体交互。本数据集旨在推动涉及人类参与的场景理解研究进展。
原文摘要 · Abstract (English)
When humans and robotic agents coexist in an environment, scene understanding becomes crucial for the agents to carry out various downstream tasks like navigation and planning. Hence, an agent must be capable of localizing and identifying actions performed by the human. Current research lacks reliable datasets for performing scene understanding within indoor environments where humans are also a part of the scene. Scene Graphs enable us to generate a structured representation of a scene or an image to perform visual scene understanding. To tackle this, we present HOIverse a synthetic dataset at the intersection of scene graph and human-object interaction, consisting of accurate and dense relationship ground truths between humans and surrounding objects along with corresponding RGB images, segmentation masks, depth images and human keypoints. We compute parametric relations between various pairs of objects and human-object pairs, resulting in an accurate and unambiguous relation definitions. In addition, we benchmark our dataset on state-of-the-art scene graph generation models to predict parametric relations and human-object interactions. Through this dataset, we aim to accelerate research in the field of scene understanding involving people.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。