arXiv:2606.19103cs.CVcs.AI2026-06

专为产品图像编辑设计数据集与训练方法,显著提升文字和品牌一致性。

ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL

论文配图:ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL
图 1 · 摘自论文原文
  • 构建87k样本SFT数据集+869张产品图像的RL数据集,专注产品特征保持。
  • 引入循环一致性奖励,使编辑后图像描述与原图高度相似,字符错误率降5倍。
  • 适用于电商、广告等需精准保留品牌和文本的视觉编辑场景。

指令驱动的图像编辑虽有进展,但在产品类场景中常无法维持品牌特征与文字准确性。现有数据集缺乏对文本保真度的明确约束,导致该能力被隐式依赖。本文提出ProductConsistency数据集,包含87,000个SFT样本、869张产品的强化学习数据,以及新的基准测试集。通过循环一致性奖励(基于原始描述与编辑后图像生成描述的相似性)指导强化学习训练。在Qwen-Image-Edit-2511和Flux.1-Kontext-dev上微调后,模型在OCR、感知评估及多模态大模型评测中均优于基线,字符错误率降低5倍,显著提升产品一致性、文字渲染与整体视觉质量。代码与流程已公开。

原文摘要 · Abstract (English)

Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions. However, in product-centric scenarios where preserving product features, branding, and textual elements are critical, current open and closed source models often struggle to maintain this fine-grained object identity. This issue is further compounded by the lack of datasets for instruction-based product image editing with text fidelity constraints, leaving it largely treated as an implicit capability of instruction-based image editing models. In this work, we introduce the ProductConsistency dataset which is designed to improve product-centric image editing. Our approach includes a supervised fine-tuning (SFT) dataset of 87k samples for product editing, a reinforcement learning (RL) dataset with 869 unique product images, and a new benchmark dataset, the ProductConsistency Benchmark, to allow rigorous and standardized evaluation of editing models. To guide RL training, we propose a Cyclic Consistency reward that enforces semantic preservation of product identity by using caption similarity between the original product description and captions generated from the edited image. We fine-tune both Qwen-Image-Edit-2511 and Flux.1-Kontext-dev using our dataset and demonstrate consistent improvements over baseline models in OCR and Perceptual metrics, and MLLM-based evaluations as well, indicating stronger product consistency, text rendering, and overall visual quality; with the Qwen-Image-Edit-2511 model achieving a 5x reduction in the character error rate. The code and pipeline is available at https://anonymous.4open.science/r/ProductConsistency-6FCC/README.md

图像编辑产品一致性强化学习OCR优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。