arXiv:2607.01222cs.CV2026-07中稿 · ECCV

用视频模型生成超复杂3D纹理,解决传统方法纹理粗糙问题。

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models

论文配图:Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models
图 1 · 摘自论文原文
  • 分步生成:先建白模,再用视频模型生成多视角外观序列。
  • 通过神经烘焙模块融合多视角信息,还原高精度纹理。
  • 适合需要精细纹理的3D内容创作者,如游戏与影视建模。

现有3D生成模型虽能生成高质量几何结构,但难以从参考图像中复现复杂纹理,主要因缺乏大规模带丰富表面外观的3D训练数据。相比之下,视觉生成模型在规模大几个数量级的数据集上训练,擅长捕捉复杂视觉模式。为此,我们提出Ink3D框架,将3D生成与大规模视频生成模型结合,实现极复杂纹理的合成。Ink3D首先使用现成的3D生成模型重建白模几何。随后,采用OrbitPainter——一种条件视频生成模型——生成密集轨道扫描视频,覆盖物体各视角外观。为将这些视角转化为一致纹理,我们引入TextureOptimizer,一个神经烘焙模块,可整合密集多视角观测,同时缓解由视频生成引起的几何不一致性。通过解耦几何与纹理生成,并利用大规模预训练视频先验,Ink3D显著提升了纹理的丰富度与真实性,优于以往方法。

原文摘要 · Abstract (English)

Recent 3D generative models can synthesize high-quality geometry but often struggle to reproduce intricate textures from reference images, largely due to the scarcity of large-scale 3D training data with rich surface appearance. In contrast, visual generative models are trained on datasets several orders of magnitude larger and excel at modeling complex visual patterns. Motivated by this gap, we introduce Ink3D, a framework that bridges 3D generation with large-scale video generative models to synthesize extremely complex textures. Ink3D first reconstructs a white-mesh geometry using an off-the-shelf 3D generation model. It then employs OrbitPainter, a conditional video generative model, to produce dense orbit-scan videos capturing object appearance across viewpoints. To convert these views into coherent textures, we introduce TextureOptimizer, a neural baking module that integrates dense multi-view observations while mitigating geometry inconsistencies arising from video generation. By decoupling geometry and texture synthesis and leveraging large-scale pretrained video priors, Ink3D enables significantly richer and more faithful texture generation than prior approaches.

3D生成纹理合成视频模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。