构建近1000张机器人采集的室内图像数据集,支持空间关系理解。
A Spatial Relationship Aware Dataset for Robotics
- 用波士顿动力Spot机器人采集带空间关系标注的图像。
- 六种场景图生成模型在该数据集上表现差异显著。
- 适合研究机器人空间推理与规划的学者使用。
真实环境中的机器人任务规划不仅需要物体识别,还需理解物体间的空间关系。我们构建了一个包含近1000张机器人采集的室内图像的数据集,图像标注了物体属性、位置及详细的空间关系。数据通过波士顿动力Spot机器人采集,并使用自研标注工具完成,涵盖相似或相同的物体以及复杂的空间布局。我们在该数据集上对六种最先进的场景图生成模型进行了基准测试,分析其推理速度和关系准确率。结果表明模型间性能差异明显,且将显式空间关系融入基础模型(如ChatGPT 4o)可显著提升其生成可执行空间感知计划的能力。数据集与标注工具已公开于https://github.com/PengPaulWang/SpatialAwareRobotDataset,支持机器人空间推理的进一步研究。
原文摘要 · Abstract (English)
Robotic task planning in real-world environments requires not only object recognition but also a nuanced understanding of spatial relationships between objects. We present a spatial-relationship-aware dataset of nearly 1,000 robot-acquired indoor images, annotated with object attributes, positions, and detailed spatial relationships. Captured using a Boston Dynamics Spot robot and labelled with a custom annotation tool, the dataset reflects complex scenarios with similar or identical objects and intricate spatial arrangements. We benchmark six state-of-the-art scene-graph generation models on this dataset, analysing their inference speed and relational accuracy. Our results highlight significant differences in model performance and demonstrate that integrating explicit spatial relationships into foundation models, such as ChatGPT 4o, substantially improves their ability to generate executable, spatially-aware plans for robotics. The dataset and annotation tool are publicly available at https://github.com/PengPaulWang/SpatialAwareRobotDataset, supporting further research in spatial reasoning for robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。