用参考图生成高保真广告图,省去繁琐调参。
RefAdGen: High-Fidelity Advertising Image Generation
- 分步设计:输入产品掩码+注意力融合模块,精准控制图像生成。
- 在10万张广告图数据集上测试,对新商品和真实场景图均保持高保真。
- 适合电商与营销领域,无需微调即可快速生成高质量广告图。
人工智能生成内容(AIGC)技术的快速发展为基于参考产品图和文本场景描述生成多样且吸引人的广告图像提供了可能,显著降低了传统营销流程中的人力成本和制作开销。然而,现有AIGC方法要么需针对每张参考图进行大量微调以实现高保真,要么难以在不同商品间保持一致的高保真度,限制了其在电商和营销领域的实际应用。为此,我们构建了大规模广告图像生成数据集AdProd-100K,其核心在于双数据增强策略,增强了对3D结构的感知能力,有助于生成更真实的图像。基于该数据集,我们提出RefAdGen框架,通过解耦设计实现高保真生成:在U-Net输入处注入产品掩码以实现精确空间控制,并采用高效的注意力融合模块(AFM)整合产品特征。该设计有效解决了现有方法中存在的保真度-效率矛盾。大量实验表明,RefAdGen达到当前最优性能,在未见商品及复杂真实场景图像上仍保持高保真与出色视觉效果,为传统工作流提供可扩展、低成本的替代方案。代码与数据集已公开于https://github.com/Anonymous-Name-139/RefAdgen。
原文摘要 · Abstract (English)
The rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques has unlocked opportunities in generating diverse and compelling advertising images based on referenced product images and textual scene descriptions. This capability substantially reduces human labor and production costs in traditional marketing workflows. However, existing AIGC techniques either demand extensive fine-tuning for each referenced image to achieve high fidelity, or they struggle to maintain fidelity across diverse products, making them impractical for e-commerce and marketing industries. To tackle this limitation, we first construct AdProd-100K, a large-scale advertising image generation dataset. A key innovation in its construction is our dual data augmentation strategy, which fosters robust, 3D-aware representations crucial for realistic and high-fidelity image synthesis. Leveraging this dataset, we propose RefAdGen, a generation framework that achieves high fidelity through a decoupled design. The framework enforces precise spatial control by injecting a product mask at the U-Net input, and employs an efficient Attention Fusion Module (AFM) to integrate product features. This design effectively resolves the fidelity-efficiency dilemma present in existing methods. Extensive experiments demonstrate that RefAdGen achieves state-of-the-art performance, showcasing robust generalization by maintaining high fidelity and remarkable visual results for both unseen products and challenging real-world, in-the-wild images. This offers a scalable and cost-effective alternative to traditional workflows. Code and datasets are publicly available at https://github.com/Anonymous-Name-139/RefAdgen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。