arXiv:2603.17735cs.CV2026-03

用3D几何约束生成一致的360度视频,自动生成高保真纹理。

TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos

  • 以3D网格为条件,通过几何特征精控视频生成过程。
  • 生成视频在多视角下保持外观一致,提升重建质量。
  • 适合需要自动化生成完整3D资产的数字内容创作者。

自动为无纹理3D模型生成逼真且自洽的外观是数字内容创作的关键挑战。大规模视频生成模型提供了一种自然方案:直接合成360度转盘视频(TTVs),不仅能作为高质量动态预览,还可作为中间表示驱动纹理合成与神经渲染。然而,现有通用视频扩散模型难以在全视角范围内保持严格的几何一致性与外观稳定性,导致其输出不适合高质量3D重建。为此,我们提出TAPESTRY框架,基于显式3D几何条件生成高保真TTVs。将3D外观生成重构为几何条件视频扩散问题:给定3D网格,先渲染并编码多模态几何特征,以像素级精度约束视频生成,从而实现高质量、一致性的TTVs生成。在此基础上,设计了下游重建方法,采用多阶段3D感知修复管道,通过旋转模型并进行上下文感知的二次生成,有效填补自遮挡区域,实现表面全覆盖。TAPESTRY生成的视频不仅是高质量动态预览,更可作为可靠的3D感知中间表示,无缝回投影至UV纹理或用于监督如3DGS等神经渲染方法,实现从无纹理网格到生产级完整3D资产的自动化生成。实验表明,本方法在视频一致性与最终重建质量上均优于现有方法。

原文摘要 · Abstract (English)

Automatically generating photorealistic and self-consistent appearances for untextured 3D models is a critical challenge in digital content creation. The advancement of large-scale video generation models offers a natural approach: directly synthesizing 360-degree turntable videos (TTVs), which can serve not only as high-quality dynamic previews but also as an intermediate representation to drive texture synthesis and neural rendering. However, existing general-purpose video diffusion models struggle to maintain strict geometric consistency and appearance stability across the full range of views, making their outputs ill-suited for high-quality 3D reconstruction. To this end, we introduce TAPESTRY, a framework for generating high-fidelity TTVs conditioned on explicit 3D geometry. We reframe the 3D appearance generation task as a geometry-conditioned video diffusion problem: given a 3D mesh, we first render and encode multi-modal geometric features to constrain the video generation process with pixel-level precision, thereby enabling the creation of high-quality and consistent TTVs. Building upon this, we also design a method for downstream reconstruction tasks from the TTV input, featuring a multi-stage pipeline with 3D-Aware Inpainting. By rotating the model and performing a context-aware secondary generation, this pipeline effectively completes self-occluded regions to achieve full surface coverage. The videos generated by TAPESTRY are not only high-quality dynamic previews but also serve as a reliable, 3D-aware intermediate representation that can be seamlessly back-projected into UV textures or used to supervise neural rendering methods like 3DGS. This enables the automated creation of production-ready, complete 3D assets from untextured meshes. Experimental results demonstrate that our method outperforms existing approaches in both video consistency and final reconstruction quality.

3D生成视频生成纹理合成几何约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。