用生成模型从物理渲染到照片级渲染,实现可控逼真图像生成。
GeRM: A Generative Rendering Model From Physically Realistic to Photorealistic

- 通过学习分布转移向量场,引导从物理渲染到照片级渲染的渐进生成。
- 在P2P-50K数据集上训练,实现文本提示驱动的局部细节增强。
- 支持多条件控制,适合影视、游戏等高保真图像生成场景。
尽管基于物理的渲染(PBR)能保证物理真实性,但实现真正照片级渲染(PRR)需要高昂的时间与人力成本,且仍难以捕捉真实世界的复杂细节。本文提出GeRM,首个从PBR到PRR(P2P)的多模态生成渲染模型。通过学习分布转移向量(DTV)场来引导生成过程,引入多条件ControlNet,结合G-buffers、文本提示和增强区域线索,逐步将PBR图像转化为PRR图像。为提升文本提示对图像分布变化的建模能力,提出残差感知转移机制,明确指定修改区域的增量更新。为监督该转换过程,构建专家指导的成对转换数据集P2P-50K,每对样本对应DTV场中的特定转移向量。大量实验表明,GeRM可生成高质量可控图像,在多种应用中优于现有最先进方法,涵盖PBR与PRR图像生成与编辑。
原文摘要 · Abstract (English)
While physically-based rendering (PBR) simulates light transport that guarantees physical realism, achieving true photorealistic rendering (PRR) demands prohibitive time and labor, and still struggles to capture the intractable richness of the real world. We propose GeRM, the first multimodal generative rendering model to bridge the gap from PBR to PRR (P2P). We formulate this P2P transition by learning a distribution transfer vector (DTV) field to direct the generative process. To achieve this, we introduce a multi-condition ControlNet that synthesizes PBR images and progressively transitions them into PRR images, guided by G-buffers, text prompts, and cues for enhanced regions. To improve the model's grasp of the image distribution shift driven by text prompts, we propose a residual perceptual transfer mechanism to associate text prompts with corresponding targeted modification regions, which more clearly defines the incremental component updates. To supervise this transfer process, we introduce a multi-agent visual language model framework to construct an expert-guided pairwise transfer dataset, named P2P-50K, where each paired sample corresponds to a specific transfer vector in the DTV field. Extensive experiments demonstrate that GeRM synthesizes high-quality controllable images and outperforms state-of-the-art baselines across diverse applications, including PBR and PRR image synthesis and editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。