arXiv:2509.05131cs.CVcs.LG2025-09被引 1

无需UV映射,单图快速生成高保真3D纹理

A Scalable Attention-Based Approach for Image-to-3D Texture Mapping

  • 用Transformer直接从单图和网格预测3D纹理场
  • 每模型生成仅需0.2秒,且比现有方法更贴近原图
  • 适合需要高效可控3D内容生成的创作者

高质量纹理对真实感3D内容创作至关重要,但现有生成方法速度慢、依赖UV映射,且常与参考图像不符。为此,我们提出一种基于Transformer的框架,直接从单张图像和网格预测3D纹理场,无需UV映射和可微渲染,实现更快纹理生成。方法结合三平面表示与基于深度的反投影损失,支持高效训练与快速推理。训练完成后,可在单次前向传播中生成高保真纹理,每形状仅需0.2秒。大量定性、定量及用户偏好评估表明,该方法在单图像纹理重建上,无论是输入图像保真度还是感知质量方面,均优于当前最优基线,凸显其在可扩展、高质量、可控3D内容创作中的实用性。

原文摘要 · Abstract (English)

High-quality textures are critical for realistic 3D content creation, yet existing generative methods are slow, rely on UV maps, and often fail to remain faithful to a reference image. To address these challenges, we propose a transformer-based framework that predicts a 3D texture field directly from a single image and a mesh, eliminating the need for UV mapping and differentiable rendering, and enabling faster texture generation. Our method integrates a triplane representation with depth-based backprojection losses, enabling efficient training and faster inference. Once trained, it generates high-fidelity textures in a single forward pass, requiring only 0.2s per shape. Extensive qualitative, quantitative, and user preference evaluations demonstrate that our method outperforms state-of-the-art baselines on single-image texture reconstruction in terms of both fidelity to the input image and perceptual quality, highlighting its practicality for scalable, high-quality, and controllable 3D content creation.

3D纹理图像生成Transformer高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。