首个面向人本放置任务的基准数据集,助力机器人理解人类意图。
Assistant Placement Aria: A Benchmark for Egocentric Placement Assistance

- 构建包含三维点云与文本描述的虚拟放置基准,支持多场景评估。
- 涵盖2D面板、就坐建议、电视放置三类任务,覆盖全局与局部约束。
- 适合研究人机协作、场景理解与智能助理的学者使用。
人类在机器人任务中涉及导航、物体操作和放置等多个方面,核心挑战在于选择符合人类意图或偏好的目标位置。本文聚焦于虚拟放置(VP)任务——基于场景上下文和以人为中心的约束,识别所有可能的目标位置。这与传统仅针对单一预定义目标位置的放置任务不同。VP问题复杂,需结合场景的几何、语义与合理性进行全局与局部推理。为此,我们提出首个专门研究该问题的基准——Assistant Placement Aria,包含合成与真实室内场景,标注了三个任务:(i) 2D面板放置,(ii) 就坐建议,(iii) 电视放置。每个场景均配有2D图像、3D点云及对象的文本描述。通过提供此基准,旨在推动这一依赖高质量数据但尚未充分探索领域的研究进展。我们还评估了多个基础模型在目标检测与分割任务上的表现。
原文摘要 · Abstract (English)
Human assistance in robotics spans around several tasks such as navigation, object manipulation, and placement, where a key challenge is selecting target destinations that align with human intentions or preferences. We focus on this challenge in the context of Virtual Placement (VP), the task of identifying all plausible target locations given scene context and human-centric constraints. This differs from traditional placement tasks that typically focus on a single, predefined target location. The VP problem is complex, as it requires both global and local reasoning about the scene's geometry, semantics, and plausibility. To address this gap, we introduce {\bf Assistant Placement Aria}, the first benchmark to explore diverse aspects of VP, including global, local, and human-centric constraints. It contains both synthetic and real indoor scenes annotated for three tasks: (i)~2D Panel Placement, (ii)~Sitting Suggestion, and (iii)~TV Placement. Each scene includes 2D images, a 3D point cloud, and a textual description of the objects within the scene. By contributing this benchmark, we aim to encourage further research in this underexplored and challenging field that is critically dependent on relevant data. We also evaluate several foundation models for object detection and segmentation on our benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。