用强化学习提升电商图像去杂质效果,更准更稳。
RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
- 用空间掩码优化+强化学习调整注意力,聚焦背景
- 复合奖励机制减少误删和视觉瑕疵,提升真实感
- 专为电商设计数据集与评测基准,适合落地应用
网络数据中,商品图片对提升用户参与度和广告效果至关重要,但水印、促销文字等干扰元素仍严重影响视觉呈现。尽管基于扩散模型的修复方法已有进展,但在商业场景下仍存在去物不可靠、领域适配有限的问题。为此,我们提出 Repainter,一种融合空间掩码轨迹优化与组相对策略优化(GRPO)的强化学习框架。该方法通过调制注意力机制强调背景上下文,生成更高奖励样本,减少错误对象插入。我们还设计了结合全局、局部与语义约束的复合奖励机制,有效降低视觉伪影和奖励欺骗问题。此外,我们构建了 EcomPaint-100K 高质量大规模电商修复数据集及标准化评测基准 EcomPaint-Bench。大量实验表明,Repainter 在复杂构图场景下显著优于现有最优方法。代码与权重将在论文接受后公开。
原文摘要 · Abstract (English)
In web data, product images are central to boosting user engagement and advertising efficacy on e-commerce platforms, yet the intrusive elements such as watermarks and promotional text remain major obstacles to delivering clear and appealing product visuals. Although diffusion-based inpainting methods have advanced, they still face challenges in commercial settings due to unreliable object removal and limited domain-specific adaptation. To tackle these challenges, we propose Repainter, a reinforcement learning framework that integrates spatial-matting trajectory refinement with Group Relative Policy Optimization (GRPO). Our approach modulates attention mechanisms to emphasize background context, generating higher-reward samples and reducing unwanted object insertion. We also introduce a composite reward mechanism that balances global, local, and semantic constraints, effectively reducing visual artifacts and reward hacking. Additionally, we contribute EcomPaint-100K, a high-quality, large-scale e-commerce inpainting dataset, and a standardized benchmark EcomPaint-Bench for fair evaluation. Extensive experiments demonstrate that Repainter significantly outperforms state-of-the-art methods, especially in challenging scenes with intricate compositions. We will release our code and weights upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。