arXiv:2607.13468cs.CVcs.LG2026-07中稿 · ICML

通过分层体素增强,从单张图像生成高分辨率3D场景

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

论文配图:HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation
图 1 · 摘自论文原文
  • 构建分层组件树,将2D图像与3D体素对齐
  • 多阶段超分辨率提升至128×128×128体素精度
  • 适合需要高质量3D重建的视觉生成任务

近期方法虽能从单张图像生成逼真的3D物体,但受限于表示分辨率,难以用于3D场景生成。本文提出HIVE-3D,一种基于分层体素增强框架的高质量3D场景生成方法。给定单张场景图像,先生成粗略初始场景,再通过图像分割和注意力检索,将2D图像组件与3D场景组件对齐。随后将这些关系组织为分层组件树,叶节点代表更细粒度的组件。最后设计体素超分辨率模型,在保持与粗体素强一致性的同时,为目标实例生成精细体素。结合该模型,对每个组件进行从粗到细的分层超分辨率处理,最终生成高分辨率、高质量的3D场景。大量实验表明,本方法显著优于先前方法,在ScanNet、Matterport3D等数据集上实现当前最优性能。

原文摘要 · Abstract (English)

Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation. In this work, we introduce HIVE-3D, a novel method for high-quality 3D scene generation based on hierarchical voxel enhancement framework. Specifically, given a single scene image as input, we first produce a coarse initial scene, then introduce image segmentation and attention-based retrieval to align 2D image components with 3D scene components. Subsequently, we organize these scene relations into a hierarchical component tree, where nodes closer to the leaves denote finer-grained components. Finally, we propose a voxel super-resolution model that generates refined voxels for the target instance while maintaining strong consistency with the coarse voxels. Equipped with this model, we perform coarse-to-fine hierarchical super-resolution on images and voxels for each component, producing a high-resolution and high-quality 3D scene. Extensive experiments demonstrate that our method significantly outperforms previous approaches, achieving state-of-the-art performance.

3D生成体素增强图像生成分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。