arXiv:2605.17102cs.GRcs.CV2026-05

用体素扩散生成无碰撞室内布局,支持高保真场景设计。

VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement

论文配图:VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement
图 1 · 摘自论文原文
  • 以锚点为条件,分步生成离散体素占据,避免空间冲突。
  • 在复杂环境中实现零碰撞布局,且保持形状多样性与物理合理性。
  • 适合需要真实感3D场景生成的设计师与机器人路径规划研究者。

我们提出VoxScene,一种专为3D场景合成设计的锚点条件体素扩散框架。现有数据驱动布局生成方法通常依赖边界框代理或隐式表示,忽视了体素结构,导致密集环境中的严重物理碰撞和结构纠缠。为此,我们转向显式、以物体为中心的体素表示,通过先前锚点和局部上下文条件,分步合成离散体素占据。利用体素的互斥性,方法消除空间模糊,确保即使在复杂环境中也实现无碰撞排列。此外,生成的高保真体素网格可作为下游资产检索的判别性几何查询。大量实验表明,该方法具有普遍性,在物理合理性方面达到当前最优水平,并显著提升形状多样性,优于现有布局规划器。

原文摘要 · Abstract (English)

We present VoxScene, a novel anchor-conditioned voxel diffusion framework tailored for 3D scene synthesis. Current data-driven layout generation techniques typically rely on bounding proxies or implicit representations, which overlook volumetric structures. This geometric blindness inevitably leads to severe physical collisions and structural entanglement, particularly in densely populated environments. To overcome these limitations, we shift the paradigm to an explicit, object-centric voxel representation. Our pipeline sequentially synthesizes discrete volumetric occupancies conditioned on prior anchors and local context. By exploiting the mutually exclusive nature of discrete voxels, our approach eliminates spatial ambiguities and guarantees collision-free arrangements, even in highly complex environments. Furthermore, the synthesized high-fidelity voxel grids serve as discriminative geometric queries for downstream asset retrieval. Extensive experiments demonstrate the universality of our method, achieving state-of-the-art physical plausibility and unlocking shape diversity compared to existing layout planners.

3D生成体素扩散场景布局

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。