arXiv:2609.06385cs.CV2026-09

无需训练即可生成大规模带纹理的3D网格,保持局部连续性和全局一致性。

Scaling 3D Generative Priors to Large-Scale Scene Meshes from Multi-View Images

论文配图:Scaling 3D Generative Priors to Large-Scale Scene Meshes from Multi-View Images
图 1 · 摘自论文原文
  • 分块生成+局部上下文增强,提升相邻区域几何连贯性。
  • 引入全局外观对齐,减少远距离区域的视觉差异。
  • 自适应分解场景,按几何复杂度动态调整分块数量。

预训练的3D生成模型虽能生成精细的几何与外观,但主要面向小范围物体生成。现有方法将大场景划分为多个小区域,并在每个区域应用预训练3D生成先验,但在扩展到大规模多视角场景时,难以保证局部几何连续性与全局外观一致性。本文提出一种无需训练的大规模纹理网格生成框架,核心思想是在提升空间细节的同时,协调局部与全局生成。引入局部上下文分块生成以增强邻近区域间的几何连续性,以及全局外观对齐机制以减少远距离区域间的外观差异。此外,通过自适应场景分解,根据输入场景几何复杂度动态确定分块数量。实验表明,该方法在几何与外观保真度上优于现有方法,支持对大规模场景进行细粒度生成。

原文摘要 · Abstract (English)

Pretrained 3D generative models produce detailed geometry and appearance but are primarily designed for object-centric generation within a limited spatial extent. Recent approaches address this limitation by partitioning large scenes into smaller spatial regions and applying pretrained 3D generative priors to each region. However, scaling tiled generation to large multi-view scenes makes it challenging to maintain local geometric continuity and global appearance consistency. We present a training-free framework for large-scale textured mesh generation from multi-view images. Our key idea is to scale tiled generation to large scenes with increased spatial detail while coordinating generation both locally and globally. We introduce local context tiled generation to improve geometric continuity between neighboring regions and global appearance alignment to reduce appearance discrepancies across distant regions. An adaptive scene decomposition further determines the number of tiles according to the input scene geometry. Experiments demonstrate improved geometric and appearance fidelity over existing approaches while enabling fine-grained generation of large-scale scenes.

3D生成纹理网格多视图重建分块生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。