用多词文本反演提升单图3D重建精度
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
- 通过多词文本反演提取图像多维度特征
- 在真实与合成数据上均超越现有方法
- 适合需要高保真3D建模的设计师与开发者
从单视角图像重建3D模型是计算机视觉中的长期难题。当前方法通常仅捕捉图像的单一关键属性(如物体类型、艺术风格),未能充分考虑形状和材质等多视角信息,且依赖神经辐射场难以还原复杂表面与纹理细节。本文提出MTFusion,结合图像与文本描述实现高保真3D重建。首先采用新型多词文本反演技术,提取涵盖图像多维特征的详细文本描述;随后利用该描述与图像生成3D模型,基于FlexiCubes框架并引入专用符号距离函数解码器,实现更快训练与更精细的表面表示。大量实验表明,MTFusion在多种合成与真实图像上均优于现有方法,消融实验证明了网络设计的有效性。
原文摘要 · Abstract (English)
Reconstructing 3D models from single-view images is a long-standing problem in computer vision. The latest advances for single-image 3D reconstruction extract a textual description from the input image and further utilize it to synthesize 3D models. However, existing methods focus on capturing a single key attribute of the image (e.g., object type, artistic style) and fail to consider the multi-perspective information required for accurate 3D reconstruction, such as object shape and material properties. Besides, the reliance on Neural Radiance Fields hinders their ability to reconstruct intricate surfaces and texture details. In this work, we propose MTFusion, which leverages both image data and textual descriptions for high-fidelity 3D reconstruction. Our approach consists of two stages. First, we adopt a novel multi-word textual inversion technique to extract a detailed text description capturing the image's characteristics. Then, we use this description and the image to generate a 3D model with FlexiCubes. Additionally, MTFusion enhances FlexiCubes by employing a special decoder network for Signed Distance Functions, leading to faster training and finer surface representation. Extensive evaluations demonstrate that our MTFusion surpasses existing image-to-3D methods on a wide range of synthetic and real-world images. Furthermore, the ablation study proves the effectiveness of our network designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。