用扩散模型生成高保真产品图像,提升电商虚拟展示效果
Preserving Product Fidelity in Large Scale Image Recontextualization with Diffusion Models
- 结合图像转视频、内外补全与负样本生成合成数据
- 在ABO和私有数据集上显著提升图像真实感与多样性
- 适合电商、虚拟展厅等需要精准产品视觉呈现的场景
我们提出一种基于文本到图像扩散模型的高保真产品图像再情境化框架,并设计了一种新颖的数据增强流水线。该流水线利用图像转视频扩散模型、内外补全及负样本生成合成训练数据,解决了真实数据收集的局限性。通过解耦产品表征并增强模型对产品特性的理解,提升了生成图像的质量与多样性。在ABO数据集和一个私有产品数据集上的评估显示,该框架能生成更逼真、更具吸引力的产品视觉呈现,适用于电子商务和虚拟产品展示等应用。
原文摘要 · Abstract (English)
We present a framework for high-fidelity product image recontextualization using text-to-image diffusion models and a novel data augmentation pipeline. This pipeline leverages image-to-video diffusion, in/outpainting & negatives to create synthetic training data, addressing limitations of real-world data collection for this task. Our method improves the quality and diversity of generated images by disentangling product representations and enhancing the model's understanding of product characteristics. Evaluation on the ABO dataset and a private product dataset, using automated metrics and human assessment, demonstrates the effectiveness of our framework in generating realistic and compelling product visualizations, with implications for applications such as e-commerce and virtual product showcasing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。