arXiv:2505.20129cs.CVcs.GR2025-05被引 17

让视觉语言模型学会理解并生成有空间结构的3D场景。

Agentic 3D Scene Generation with Spatially Contextualized VLMs

论文配图:Agentic 3D Scene Generation with Spatially Contextualized VLMs
图 1 · 摘自论文原文
  • 用动态空间上下文增强VLM,使其具备3D结构化推理能力。
  • 可生成高质量3D场景,支持自动验证与人体工学调整。
  • 适合开发智能交互、沉浸式模拟等空间感知应用。

尽管视觉语言模型(VLM)在多模态内容生成方面取得进展,但其对结构化3D场景的理解与生成能力仍受限,制约了其在具身AI、沉浸式仿真和交互式3D应用中的使用。本文提出一种新范式,通过持续演化的空间上下文,使VLM能够生成、理解与编辑复杂3D环境。该上下文由三部分构成:提供高层语义蓝图的场景画像、捕捉物体级几何的语义点云,以及编码丰富空间关系(包括一元、二元及高阶约束)的场景超图。三者共同构建了一个结构化、几何感知的工作记忆,融合VLM的多模态推理与3D结构理解能力,实现有效空间推理。基于此,我们构建了代理式3D场景生成流水线,使VLM能迭代读取并更新空间上下文。流程包含高质量资产生成与几何修复、环境自动设置与验证,以及由场景超图引导的人体工学调整。实验表明,该框架可处理多样且复杂的输入,展现出先前工作未见的泛化能力。进一步结果表明,注入空间上下文后,VLM能执行交互式场景编辑与路径规划等下游任务,显示出在计算机图形学、3D视觉和具身应用中构建空间智能系统的重要潜力。

原文摘要 · Abstract (English)

Despite recent advances in multimodal content generation enabled by vision-language models (VLMs), their ability to reason about and generate structured 3D scenes remains largely underexplored. This limitation constrains their utility in spatially grounded tasks such as embodied AI, immersive simulations, and interactive 3D applications. We introduce a new paradigm that enables VLMs to generate, understand, and edit complex 3D environments by injecting a continually evolving spatial context. Constructed from multimodal input, this context consists of three components: a scene portrait that provides a high-level semantic blueprint, a semantically labeled point cloud capturing object-level geometry, and a scene hypergraph that encodes rich spatial relationships, including unary, binary, and higher-order constraints. Together, these components provide the VLM with a structured, geometry-aware working memory that integrates its inherent multimodal reasoning capabilities with structured 3D understanding for effective spatial reasoning. Building on this foundation, we develop an agentic 3D scene generation pipeline in which the VLM iteratively reads from and updates the spatial context. The pipeline features high-quality asset generation with geometric restoration, environment setup with automatic verification, and ergonomic adjustment guided by the scene hypergraph. Experiments show that our framework can handle diverse and challenging inputs, achieving a level of generalization not observed in prior work. Further results demonstrate that injecting spatial context enables VLMs to perform downstream tasks such as interactive scene editing and path planning, suggesting strong potential for spatially intelligent systems in computer graphics, 3D vision, and embodied applications. Project page: https://spatctxvlm.github.io/project_page/.

3D生成视觉语言模型空间推理具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。