提出细粒度约束评估框架,精准检验文本生成3D场景的准确性
Apples on the Table? Evaluating Text-Guided 3D Scene Synthesis via Fine-Grained Constraint Verification
- 将文本描述拆解为原子级约束,逐条验证物体位置关系
- 现有方法在新评测中成功率不足10%,暴露严重语义错位
- 适合关注3D生成细节对齐与评估标准的研究者
从用户文本描述中准确合成3D场景对发展具身智能体至关重要。现有评估方法要么仅捕捉合成场景与描述的粗粒度相似性,要么忽略物体空间布局的合理性。本文提出LEGO基准数据集,将每个用户描述与人工标注的约束及参考场景配对,并设计LEGO-Eval评估框架,将描述分解为基本约束,利用工具将文本指代映射到3D物体并推理其空间关系。实验表明:(i) LEGO-Eval比现有方法更准确识别语义错位;(ii) 当前主流场景生成方法在LEGO-Eval上成功率最高仅为10%。
原文摘要 · Abstract (English)
Accurately synthesizing 3D scenes from user-provided text descriptions is crucial for developing embodied agents. Despite the importance of scene-description alignment, existing evaluation methods for such text-guided 3D scene synthesis either capture only coarse similarity between the synthesized scene and the user description, or ignore the spatial reasoning for verifying object placement. None of them addressed the fine-grained constraints (e.g., X needs to be in the scene in a Y manner) implied by the description from users. To address this, we introduce LEGO, a benchmark dataset that pairs each user description with human-annotated constraints and a reference scene, and LEGO-Eval, an evaluation framework that decomposes a description into atomic constraints and verifies each one using tools that ground textual references to 3D objects and reason about their spatial relationships. We show that (i) LEGO-Eval evaluates misalignment far more accurately than existing methods and (ii) current scene synthesis approaches achieve at most 10% success rate in LEGO-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。