arXiv:2412.02193cs.CVcs.AI2024-12CVPR被引 102

用视觉语言模型优化3D布局,让物体摆放更符合语义指令。

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

  • 通过视觉语言模型生成双向增强的场景表示
  • 实现可微分优化,确保布局物理合理
  • 适合需要精准空间规划的应用场景

空间推理是人类认知的核心能力,使人能够直观理解并操作三维空间中的物体。尽管基础模型在某些基准测试中表现优异,但在根据开放式语言指令在密集且物理受限环境中排列物体等3D推理任务上仍存在困难。我们提出LayoutVLM,一种利用视觉语言模型(VLMs)语义知识的框架与场景布局表示,并支持可微分优化以保证物理合理性。LayoutVLM通过视觉标记图像生成两个相互强化的表示,并采用自洽解码过程提升VLM的空间规划能力。实验表明,LayoutVLM克服了现有大语言模型和基于约束方法的局限性,生成的3D布局更符合输入语言指令的语义意图,且在现有场景数据集上提取布局表示微调VLMs,可有效提升其推理性能。

原文摘要 · Abstract (English)

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still struggle with 3D reasoning tasks like arranging objects in space according to open-ended language instructions, particularly in dense and physically constrained environments. We introduce LayoutVLM, a framework and scene layout representation that exploits the semantic knowledge of Vision-Language Models (VLMs) and supports differentiable optimization to ensure physical plausibility. LayoutVLM employs VLMs to generate two mutually reinforcing representations from visually marked images, and a self-consistent decoding process to improve VLMs spatial planning. Our experiments show that LayoutVLM addresses the limitations of existing LLM and constraint-based approaches, producing physically plausible 3D layouts better aligned with the semantic intent of input language instructions. We also demonstrate that fine-tuning VLMs with the proposed scene layout representation extracted from existing scene datasets can improve their reasoning performance.

3D布局视觉语言模型可微分优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。