arXiv:2512.08535cs.CV2025-12

用AI生成图像增强3D模型细节,让虚拟物体更逼真。

Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement

  • 用GPT-4o生成多视角图像,对齐3D结构提升细节。
  • 在多个3D生成框架上实现顶尖的逼真度表现。
  • 适合做高质量3D内容生成的研究者与开发者。

尽管近期的3D原生生成器在几何可靠性方面取得显著进展,但在外观真实感方面仍存在不足。主要障碍在于缺乏多样且高质量的真实世界3D资产,尤其缺少丰富的纹理细节,因为场景尺度多样、物体非刚性运动以及3D扫描仪精度有限,使得数据采集极具挑战。为此,我们提出Photo3D框架,利用GPT-4o-Image模型生成的图像驱动高保真3D生成。考虑到生成图像可能因缺乏多视角一致性而扭曲3D结构,我们设计了结构对齐的多视角合成流程,并构建了一个与3D几何配对的细节增强多视角数据集。在此基础上,提出一种真实感细节增强方案,通过感知特征适配和语义结构匹配,在保持与3D原生几何结构一致的同时,实现外观一致性与真实细节。该方案可通用适配多种3D原生生成器,并提供针对性训练策略,以优化几何-纹理耦合与解耦生成范式。实验表明,Photo3D在多种3D生成范式下均表现优异,达到当前最先进的逼真3D生成水平。

原文摘要 · Abstract (English)

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-world 3D assets with rich texture details, since capturing such data is intrinsically difficult due to the diverse scales of scenes, non-rigid motions of objects, and the limited precision of 3D scanners. We introduce Photo3D, a framework for advancing photorealistic 3D generation, which is driven by the image data generated by the GPT-4o-Image model. Considering that the generated images can distort 3D structures due to their lack of multi-view consistency, we design a structure-aligned multi-view synthesis pipeline and construct a detail-enhanced multi-view dataset paired with 3D geometry. Building on it, we present a realistic detail enhancement scheme that leverages perceptual feature adaptation and semantic structure matching to enforce appearance consistency with realistic details while preserving the structural consistency with the 3D-native geometry. Our scheme is general to different 3D-native generators, and we present dedicated training strategies to facilitate the optimization of geometry-texture coupled and decoupled 3D-native generation paradigms. Experiments demonstrate that Photo3D generalizes well across diverse 3D-native generation paradigms and achieves state-of-the-art photorealistic 3D generation performance.

3D生成图像生成细节增强GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。