arXiv:2605.13852cs.GRcs.CV2026-05中稿 · CVPR被引 1

让3D生成更真实,通过分离控制信号与视觉风格。

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning

论文配图:Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning
图 1 · 摘自论文原文
  • 用残差适配器显式分离视觉域,避免模型绑定合成风格。
  • 在文本到多视角生成中,真实感显著提升,3D一致性保持良好。
  • 适合需要高真实感3D生成的研究者和工业应用开发者。

我们常希望生成既逼真又3D一致的图像,具备精确的几何、材质和视角控制。传统方法通过微调预训练于数十亿张真实图像的图像生成器,并使用合成3D资产的渲染图进行训练,虽能学习控制信号,但因照片与渲染图之间的域差距,常导致图像失真。我们发现,问题主要源于模型将控制信号的存在与合成外观错误关联。为此,提出Realiz3D——一种轻量级扩散模型训练框架,通过引入协变量驱动的小型残差适配器,显式分离视觉域(真实或合成)与其他控制信号。该设计使生成器在获得可控性的同时,不依赖特定视觉域。通过分析扩散模型中不同层与去噪步骤的作用,优化训练与推理策略,进一步缩小域间差距。实验证明,Realiz3D在文本到多视角生成和基于3D输入的纹理生成任务中,可生成3D一致且高度逼真的图像。

原文摘要 · Abstract (English)

We often aim to generate images that are both photorealistic and 3D-consistent, adhering to precise geometry, material, and viewpoint controls. Typically, this is achieved by fine-tuning an image generator, pre-trained on billions of real images, using renders of synthetic 3D assets, where annotations for control signals are available. While this approach can learn the desired controls, it often compromises the realism of the images due to domain gap between photographs and renders. We observe that this issue largely arises from the model learning an unintended association between the presence of control signals and the synthetic appearance of the images. To address this, we introduce Realiz3D, a lightweight framework for training diffusion models, that decouples controls and visual domain. The key idea is to explicitly learn visual domain, real or synthetic, separately from other control signals by introducing a co-variate that, fed into small residual adapters, shifts the domain. Then, the generator can be trained to gain controllability, without fitting to specific visual domain. In this way, the model can be guided to produce realistic images even when controls are applied. We enhance control transferability to the real domain by leveraging insights about roles of different layers and denoising steps in diffusion-based generators, informing new training and inference strategies that further mitigate the gap. We demonstrate the advantages of Realiz3D in tasks as text-to-multiview generation and texturing from 3D inputs, producing outputs that are 3D-consistent and photorealistic.

3D生成扩散模型真实感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。