arXiv:2505.05288cs.CVcs.AI2025-05ICCV被引 10

根据文字描述在真实3D场景中自动放置物体,支持智能交互与场景理解。

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

  • 通过文本提示指导3D物体在真实场景中的合理位置摆放。
  • 提出首个针对该任务的基准数据集与评估协议,支持多解合理性验证。
  • 适用于3D大模型训练与评估,推动通用三维语言模型发展。

我们提出全新的任务——在真实3D场景中进行语言引导的物体放置。给定一个3D场景的点云、一个3D资产和一段描述其应放置位置的文本提示,目标是找到符合提示且在几何上合理的放置位置。与现有3D场景中的语言定位任务(如物体接地)相比,该任务具有多重有效解、需理解3D空间关系与自由空间等挑战。为此,我们首次建立该任务的基准测试集与评估协议,并构建用于训练3D大模型的新数据集,提出首个非平凡基线方法。我们认为该任务及新基准有望成为评估和比较通用3D大模型的重要组成部分。

原文摘要 · Abstract (English)

We introduce the novel task of Language-Guided Object Placement in Real 3D Scenes. Our model is given a 3D scene's point cloud, a 3D asset, and a textual prompt broadly describing where the 3D asset should be placed. The task here is to find a valid placement for the 3D asset that respects the prompt. Compared with other language-guided localization tasks in 3D scenes such as grounding, this task has specific challenges: it is ambiguous because it has multiple valid solutions, and it requires reasoning about 3D geometric relationships and free space. We inaugurate this task by proposing a new benchmark and evaluation protocol. We also introduce a new dataset for training 3D LLMs on this task, as well as the first method to serve as a non-trivial baseline. We believe that this challenging task and our new benchmark could become part of the suite of benchmarks used to evaluate and compare generalist 3D LLM models.

3D生成语言引导场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。