用视频模型理解3D场景,让生成更真实一致。
Video Perception Models for 3D Scene Synthesis
- 融合视频模型的物理常识,统一分析物体语义与位置
- 支持图文提示,跨视角物体布局保持一致
- 引入第一人称视角评分,评估场景合理性与连贯性
传统3D场景合成依赖专家知识和大量人工操作。自动化可显著推动建筑设计、机器人仿真、虚拟现实和游戏等领域。现有方法多依赖大语言模型(LLMs)的常识推理或图像生成模型的强视觉先验,但当前LLMs在3D空间推理能力有限,难以生成真实连贯的3D场景;而基于图像生成的方法常受限于视角选择和多视图不一致。本文提出视频感知3D场景合成框架VIPScene,利用视频生成模型中编码的3D物理世界常识,确保场景布局一致性和物体跨视角放置合理。VIPScene接受文本和图像提示,融合视频生成、前馈3D重建与开放词汇感知模型,对场景中每个物体进行语义与几何分析,实现高真实感且结构一致的灵活场景生成。为提升分析精度,我们进一步提出第一人称视角评分(FPVScore),通过连续第一人称视角充分利用多模态大语言模型的推理能力进行一致性与合理性评估。大量实验表明,VIPScene显著优于现有方法,并在多样化场景中具有良好泛化性。代码将公开。
原文摘要 · Abstract (English)
Traditionally, 3D scene synthesis requires expert knowledge and significant manual effort. Automating this process could greatly benefit fields such as architectural design, robotics simulation, virtual reality, and gaming. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or strong visual priors of modern image generation models. However, current LLMs demonstrate limited 3D spatial reasoning ability, which restricts their ability to generate realistic and coherent 3D scenes. Meanwhile, image generation-based methods often suffer from constraints in viewpoint selection and multi-view inconsistencies. In this work, we present Video Perception models for 3D Scene synthesis (VIPScene), a novel framework that exploits the encoded commonsense knowledge of the 3D physical world in video generation models to ensure coherent scene layouts and consistent object placements across views. VIPScene accepts both text and image prompts and seamlessly integrates video generation, feedforward 3D reconstruction, and open-vocabulary perception models to semantically and geometrically analyze each object in a scene. This enables flexible scene synthesis with high realism and structural consistency. For more precise analysis, we further introduce First-Person View Score (FPVScore) for coherence and plausibility evaluation, utilizing continuous first-person perspective to capitalize on the reasoning ability of multimodal large language models. Extensive experiments show that VIPScene significantly outperforms existing methods and generalizes well across diverse scenarios. The code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。