针对多参考图生图难题,提出动态奖励优化框架,显著提升生成质量。
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

- 分两阶段训练:先微调基础能力,再用动态奖励重加权优化
- 在混合类型参考图增多时,性能下降明显,但新方法有效缓解
- 适合需要高质量多参考图像生成的研究者与开发者
尽管个性化图像生成已取得显著进展,多参考图像生成(MRIG)仍具挑战性。现有基准无法充分评估复杂MRIG场景,制约该领域发展。为此,我们提出OmniRef-Bench,涵盖多种参考图像类型组合及大量参考图像,更全面评估模型表现。实验表明,主流开源模型在复杂MRIG任务中表现不佳,且随着混合类型参考图像数量增加,性能显著下降。为解决此问题,我们提出DyRef,一种两阶段训练框架:第一阶段通过监督微调赋予模型处理复杂任务的基础能力;第二阶段引入难度感知优势重加权(DAR)和判别性奖励缩放(DRS),前者动态调整优化目标以应对大量混合类型参考图,后者放大组内奖励差异以实现更有效的策略优化。实验显示,DyRef显著提升开源模型在OmniRef-Bench和单图编辑基准上的表现,验证了方法的有效性与泛化能力。
原文摘要 · Abstract (English)
While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. Most existing benchmarks fail to adequately evaluate complex MRIG scenarios, hindering further progress in this area. To better assess model performance on complex MRIG tasks, we introduce OmniRef-Bench, a benchmark that covers complex combinations of reference image types and a large number of reference images. Evaluations on OmniRef-Bench show that mainstream open-source models struggle in complex MRIG scenarios, and their performance deteriorates significantly as the number of mixed-type reference images increases. To address this issue, we propose DyRef, a two-stage training framework. In the first stage, supervised fine-tuning equips the model with the basic capability to handle complex MRIG tasks. In the second stage, we introduce Difficulty-aware Advantage Reweighting (DAR) and Discriminative Reward Scaling (DRS). DAR dynamically adjusts the optimization objective to improve performance when handling a large number of mixed-type reference images. DRS enlarges intra-group reward differences for more effective policy optimization. Experiments demonstrate that DyRef significantly improves the performance of open-source models on OmniRef-Bench and single-image editing benchmarks, demonstrating the effectiveness and generalization capability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。