用多维度评估验证奖励提升多参考图像编辑的一致性。
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

- 设计多维度评估验证奖励,分解视觉标准并验证每项判断
- 在Qwen-Image-Edit基础上提升一致性与整体和谐度,媲美NanoBanana
- 无需修改架构,适配现成编辑器,适合追求高一致性的图像编辑研究者
尽管近期图像编辑模型进展迅速,多参考编辑仍面临保持参考间视觉一致性与整体和谐性的挑战。强化学习在文本到图像生成和单图编辑中表现优异,但扩展至多参考编辑受限于缺乏能捕捉多图关系约束的奖励模型。直接使用多模态大语言模型(MLLM)作为零样本评估器,又面临长篇推理易幻觉与短判别推理能力弱的矛盾。为此,我们提出多维度评估验证奖励(EVR)。EVR将评估分解为多个视觉标准;对每项标准,由MLLM评估器生成多个候选假设,再由验证器基于具体视觉证据判断其真伪,从而生成可靠且细粒度的奖励信号。结合可扩展的数据流水线,本方法可在不改变模型架构的前提下,对现成编辑器进行强化学习微调。大量实验表明,相较基线Qwen-Image-Edit,性能显著提升,一致性与和谐度达到甚至超过NanoBanana水平。
原文摘要 · Abstract (English)
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。