构建3D场景重排的语言引导基准,测试模型理解空间指代的能力
ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

- 基于CLEVR构建33826条指令,涵盖多层指代关系与放置场景
- 模型在满足语义关系上表现优于物理合理性,最近和介于之间最难
- 适合研究视觉语言模型在三维空间推理与场景理解方向的开发者
我们提出ReRef-3D,一个用于语言引导3D场景重排的基准数据集。包含998个源自CLEVR的场景,共33,826条指令,覆盖16类放置任务及直接、单跳、双跳指代。每条指令需生成有效的新放置位置,而非单一坐标。评估时将预测位置插入场景,重新计算空间关系并验证语义匹配与物理可行性。每条指令还附有经验证的自然化重写版本。微调后,LLaVA-3D、3D-LLM与PlaceIt3D分别在68.3%、31.6%和22.4%的指令中产生有效放置。总体上,语义关系满足率高于物理合理性;最近和介于关系最难实现;表述方式对性能影响较小。
原文摘要 · Abstract (English)
We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each instruction must be resolved into a valid new placement position. Given that an instruction defines a region of acceptable placements rather than one coordinate, our evaluation inserts a prediction into the scene, recomputes relations, and tests relation satisfaction and physical validity. Each instruction also includes a verified naturalized rewrite. After fine-tuning, LLaVA-3D, 3D-LLM, and PlaceIt3D produce valid placements for 68.3%, 31.6%, and 22.4% of instructions, respectively. Across models, relation satisfaction surpasses physical validity, relations such as nearest and between are the most difficult, and phrasing has minimal effect on performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。