用参考图生成高保真人货图像,细节更清晰。
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
- 引入共享增强注意力机制,聚焦产品细节特征
- 设计高频监督损失,实现像素级精准指导
- 构建4万样本新数据集,提升生成真实感
人货图像在广告、电商和数字营销中至关重要,其核心挑战在于保持产品细节的高保真度。现有基于参考图的修复方法存在三大问题:缺乏多样化的大型训练数据、模型难聚焦产品细节保留、粗粒度监督难以实现精确引导。为此,我们提出HiFi-Inpaint框架,通过引入共享增强注意力(SEA)模块优化细粒度产品特征,并设计细节感知损失(DAL),利用高频图实现像素级精确监督。此外,我们构建了新数据集HP-Image-40K,基于自合成数据并经自动过滤处理。实验表明,该方法达到当前最优性能,能生成细节丰富的人货图像。
原文摘要 · Abstract (English)
Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. The essential challenge of generating such images lies in ensuring the high-fidelity preservation of product details. Among existing paradigms, reference-based inpainting offers a targeted solution by leveraging product reference images to guide the inpainting process. However, limitations remain in three key aspects: the lack of diverse large-scale training data, the struggle of current models to focus on product detail preservation, and the inability of coarse supervision for achieving precise guidance. To address these issues, we propose HiFi-Inpaint, a novel high-fidelity reference-based inpainting framework tailored for generating human-product images. HiFi-Inpaint introduces Shared Enhancement Attention (SEA) to refine fine-grained product features and Detail-Aware Loss (DAL) to enforce precise pixel-level supervision using high-frequency maps. Additionally, we construct a new dataset, HP-Image-40K, with samples curated from self-synthesis data and processed with automatic filtering. Experimental results show that HiFi-Inpaint achieves state-of-the-art performance, delivering detail-preserving human-product images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。