arXiv:2508.02155cs.CVcs.AI2025-08被引 4

用参考图+文字精准生成电商背景,保持商品一致性。

DreamPainter: Image Background Inpainting for E-commerce Scenarios

  • 结合文本提示与参考图像双重控制生成
  • 在40万电商数据上训练,显著提升商品一致性
  • 适合需要高保真商品背景的电商平台

尽管基于扩散模型的图像生成已广泛应用,电商场景中的背景生成仍面临挑战。首要挑战是确保生成商品与输入产品一致,同时保持合理的空间布局、和谐的阴影与反光;现有修复方法因缺乏领域专用数据而难以解决。第二个挑战在于仅依赖文本提示进行图像控制,难以有效融合视觉信息实现精确控制。为此,我们构建了DreamEcom-400K,一个高质量电商数据集,包含准确的产品实例掩码、背景参考图、文本提示及美学良好的商品图像。基于该数据集,我们提出DreamPainter框架,不仅使用文本提示控制,还灵活引入参考图像作为额外控制信号。大量实验表明,该方法显著优于现有最先进方法,在保持高商品一致性的同时,有效融合文本提示与参考图像信息。

原文摘要 · Abstract (English)

Although diffusion-based image genenation has been widely explored and applied, background generation tasks in e-commerce scenarios still face significant challenges. The first challenge is to ensure that the generated products are consistent with the given product inputs while maintaining a reasonable spatial arrangement, harmonious shadows, and reflections between foreground products and backgrounds. Existing inpainting methods fail to address this due to the lack of domain-specific data. The second challenge involves the limitation of relying solely on text prompts for image control, as effective integrating visual information to achieve precise control in inpainting tasks remains underexplored. To address these challenges, we introduce DreamEcom-400K, a high-quality e-commerce dataset containing accurate product instance masks, background reference images, text prompts, and aesthetically pleasing product images. Based on this dataset, we propose DreamPainter, a novel framework that not only utilizes text prompts for control but also flexibly incorporates reference image information as an additional control signal. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, maintaining high product consistency while effectively integrating both text prompt and reference image information.

图像修复电商生成扩散模型多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。