用等轴视图构建分层3D场景,实现自然布局与交互编辑。
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
- 将房间视为可分解的层级对象,支持细粒度控制。
- 视频扩散+形态补全,解决遮挡与阴影问题,生成完整物体。
- 适合需要真实感3D场景的交互应用,如游戏或虚拟设计。
场景级3D生成是多媒体与计算机图形学的重要前沿,但现有方法或类别有限,或缺乏交互编辑灵活性。本文提出HiScene,一种新型分层框架,弥合2D图像生成与3D物体生成之间的差距,生成具有组合身份与美学内容的高保真场景。核心思想是将场景视为等轴视图下的层级“对象”,房间作为复杂对象可进一步分解为可操作的物品。该分层结构使3D内容与2D表示对齐,同时保持组合结构。为确保每个分解实例的完整性与空间对齐,我们提出基于视频扩散的形态补全技术,有效处理物体间的遮挡与阴影,并引入形状先验以保证场景内空间一致性。实验表明,该方法生成的物体布局更自然,实例更完整,适用于交互应用,且保持物理合理性与用户输入的一致性。
原文摘要 · Abstract (English)
Scene-level 3D generation represents a critical frontier in multimedia and computer graphics, yet existing approaches either suffer from limited object categories or lack editing flexibility for interactive applications. In this paper, we present HiScene, a novel hierarchical framework that bridges the gap between 2D image generation and 3D object generation and delivers high-fidelity scenes with compositional identities and aesthetic scene content. Our key insight is treating scenes as hierarchical "objects" under isometric views, where a room functions as a complex object that can be further decomposed into manipulatable items. This hierarchical approach enables us to generate 3D content that aligns with 2D representations while maintaining compositional structure. To ensure completeness and spatial alignment of each decomposed instance, we develop a video-diffusion-based amodal completion technique that effectively handles occlusions and shadows between objects, and introduce shape prior injection to ensure spatial coherence within the scene. Experimental results demonstrate that our method produces more natural object arrangements and complete object instances suitable for interactive applications, while maintaining physical plausibility and alignment with user inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。