arXiv:2503.04919cs.CV2025-03CVPR被引 25

用几何推理增强大模型对物体摆放的常识理解。

FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement

  • 结合大模型语义能力与几何约束求解,实现精准物体定位。
  • 在复杂场景中摆放效果优于现有方法,兼顾几何合理与常识符合。
  • 适合需要高精度3D场景生成的研究与应用者。

3D资产的场景生成面临双重挑战:既要具备高层次语义理解,又需低层次几何推理能力。虽然多模态大语言模型(MLLMs)在语义任务上表现优异,但在3D场景生成中受限于对3D几何的弱定位能力。本文提出新框架FirePlace,利用现有MLLMs完成三项任务:(1) 进行3D几何推理并提取场景中的几何细节;(2) 构建并求解提取出的底层几何约束;(3) 通过剪枝保留符合常识的最终摆放位置。通过融合几何推理与真实世界的常识理解,该方法可生成既满足几何约束又符合高层语义常识的物体放置方案。实验表明,在具有复杂几何结构的场景中,该方法能更有效地进行物体摆放,性能超越先前工作。

原文摘要 · Abstract (English)

Scene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Models (MLLMs) excel at semantic tasks, their application to 3D scene generation is hindered by their limited grounding on 3D geometry. In this paper, we investigate how to best work with MLLMs in an object placement task. Towards this goal, we introduce a novel framework, FirePlace, that applies existing MLLMs in (1) 3D geometric reasoning and the extraction of relevant geometric details from the 3D scene, (2) constructing and solving geometric constraints on the extracted low-level geometry, and (3) pruning for final placements that conform to common sense. By combining geometric reasoning with real-world understanding of MLLMs, our method can propose object placements that satisfy both geometric constraints as well as high-level semantic common-sense considerations. Our experiments show that these capabilities allow our method to place objects more effectively in complex scenes with intricate geometry, surpassing the quality of prior work.

3D生成几何推理常识理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。