用扩散模型将渲染人像转为真实感图像,保留衣物纹理细节。
FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models
- 通过注入真实与渲染图像知识,增强扩散模型控制力。
- 生成图像在视觉质量上显著优于现有方法,纹理保留更精细。
- 适合需要高保真服装图像的虚拟试衣、数字人生成场景。
建模和生成逼真人像一直是多领域研究重点,其难点在于人体高度可动且结构复杂。渲染算法虽能模拟成像过程,但受限于模型变量精度与计算效率;生成模型虽能产出生动人像,却缺乏可控性与可编辑性。本文研究如何提升渲染图像的真实感,利用扩散模型在渲染基础上实现可控生成。提出两阶段框架:域知识注入(DKI)与真实图像生成(RIG)。DKI中,采用真实图像正样本微调与渲染图像负样本嵌入,向预训练文本到图像扩散模型注入领域知识。RIG阶段,基于输入渲染图生成真实对应图像,引入纹理保持注意力控制(TAC),通过UNet结构解耦特征保留细粒度服装纹理。同时构建了SynFashion数据集,包含多样纹理的高质量数字服装图像。大量实验表明,该方法在渲染到真实图像转换任务中具有显著优势与有效性。
原文摘要 · Abstract (English)
Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accuracy of modeled variables and the efficiency of computation. Generative models can produce impressively vivid human images, however still lacking in controllability and editability. This paper studies photorealism enhancement of rendered images, leveraging generative power from diffusion models on the controlled basis of rendering. We introduce a novel framework to translate rendered images into their realistic counterparts, which consists of two stages: Domain Knowledge Injection (DKI) and Realistic Image Generation (RIG). In DKI, we adopt positive (real) domain finetuning and negative (rendered) domain embedding to inject knowledge into a pretrained Text-to-image (T2I) diffusion model. In RIG, we generate the realistic image corresponding to the input rendered image, with a Texture-preserving Attention Control (TAC) to preserve fine-grained clothing textures, exploiting the decoupled features encoded in the UNet structure. Additionally, we introduce SynFashion dataset, featuring high-quality digital clothing images with diverse textures. Extensive experimental results demonstrate the superiority and effectiveness of our method in rendered-to-real image translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。