arXiv:2503.03664cs.CVcs.AI2025-03被引 4

用文字生成高保真3D模型,全流程自动化。

A Generative Approach to High Fidelity 3D Reconstruction from Text Data

  • 文本转图像后,经去反光、超分辨率等处理提升质量
  • 通过深度学习将2D图像转为带几何细节的3D体素模型
  • 适合做AR/VR内容生成,对语义与结构都要求高

生成式AI与计算机视觉的结合,推动了从文本描述生成三维表示的突破。本研究提出一个全自动流程,整合文本到图像生成、图像处理及深度学习去反光与3D重建技术。利用Stable Diffusion等先进生成模型,将自然语言提示转化为高质量图像,再由强化学习代理增强并使用Stable Delight模型去除反光。随后应用高级图像超分与背景移除技术提升视觉保真度。这些优化后的二维图像通过复杂机器学习算法转化为体素3D模型,准确捕捉空间关系与几何特征。该方法有效解决语义一致性、几何复杂性与细节保留等挑战。实验将评估不同领域与复杂度下的重建质量、语义准确性和几何保真度。研究成果对增强现实(AR)、虚拟现实(VR)和数字内容创作具有重要意义。

原文摘要 · Abstract (English)

The convergence of generative artificial intelligence and advanced computer vision technologies introduces a groundbreaking approach to transforming textual descriptions into three-dimensional representations. This research proposes a fully automated pipeline that seamlessly integrates text-to-image generation, various image processing techniques, and deep learning methods for reflection removal and 3D reconstruction. By leveraging state-of-the-art generative models like Stable Diffusion, the methodology translates natural language inputs into detailed 3D models through a multi-stage workflow. The reconstruction process begins with the generation of high-quality images from textual prompts, followed by enhancement by a reinforcement learning agent and reflection removal using the Stable Delight model. Advanced image upscaling and background removal techniques are then applied to further enhance visual fidelity. These refined two-dimensional representations are subsequently transformed into volumetric 3D models using sophisticated machine learning algorithms, capturing intricate spatial relationships and geometric characteristics. This process achieves a highly structured and detailed output, ensuring that the final 3D models reflect both semantic accuracy and geometric precision. This approach addresses key challenges in generative reconstruction, such as maintaining semantic coherence, managing geometric complexity, and preserving detailed visual information. Comprehensive experimental evaluations will assess reconstruction quality, semantic accuracy, and geometric fidelity across diverse domains and varying levels of complexity. By demonstrating the potential of AI-driven 3D reconstruction techniques, this research offers significant implications for fields such as augmented reality (AR), virtual reality (VR), and digital content creation.

3D生成文本生成扩散模型AR/VR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。