arXiv:2504.10466cs.CV2025-04被引 1

无需训练,将扁平手绘图一键生成3D模型。

Art3D: Training-Free 3D Generation from Flat-Colored Illustration

  • 利用预训练模型提取结构语义特征,增强2D图的立体感。
  • 在100+张扁平色图像上验证,生成3D效果更真实稳定。
  • 适合艺术创作、游戏设计等需要快速转3D的场景。

大规模预训练的图像到3D生成模型在多样形状生成方面表现卓越。然而,当参考图像为手绘风格的扁平着色图时,由于缺乏三维视觉暗示,大多数模型难以生成合理3D资产,而这类输入正是艺术创作中最易用的模态。为此,我们提出Art3D,一种无需训练的方法,可将扁平着色的2D设计提升为3D模型。通过结合预训练2D图像生成模型的结构与语义特征,以及基于视觉语言模型(VLM)的真实感评估,Art3D有效增强了参考图像的三维视觉暗示,从而简化了从2D到3D的生成流程,并展现出对多种绘画风格的适应性。为评估现有图像到3D模型在无三维感的扁平色图像上的泛化能力,我们收集了一个包含超过100个样本的新数据集Flat-2D。实验结果表明,Art3D具备优异的泛化性能和良好的实用性。代码与数据集将公开于项目主页:https://joy-jy11.github.io/。

原文摘要 · Abstract (English)

Large-scale pre-trained image-to-3D generative models have exhibited remarkable capabilities in diverse shape generations. However, most of them struggle to synthesize plausible 3D assets when the reference image is flat-colored like hand drawings due to the lack of 3D illusion, which are often the most user-friendly input modalities in art content creation. To this end, we propose Art3D, a training-free method that can lift flat-colored 2D designs into 3D. By leveraging structural and semantic features with pre-trained 2D image generation models and a VLM-based realism evaluation, Art3D successfully enhances the three-dimensional illusion in reference images, thus simplifying the process of generating 3D from 2D, and proves adaptable to a wide range of painting styles. To benchmark the generalization performance of existing image-to-3D models on flat-colored images without 3D feeling, we collect a new dataset, Flat-2D, with over 100 samples. Experimental results demonstrate the performance and robustness of Art3D, exhibiting superior generalizable capacity and promising practical applicability. Our source code and dataset will be publicly available on our project page: https://joy-jy11.github.io/ .

3D生成图像转3D手绘转3D

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。