arXiv:2512.04784cs.CV2025-12被引 9

用强化学习提升图像生成的一致性,让角色风格和逻辑连贯更自然。

PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling

  • 通过成对评分模型学习视觉一致性,无需大量标注数据。
  • 在两个子任务上超越现有方法,训练效率与稳定性显著提升。
  • 适合需要角色连贯性的故事插画、动画设计等场景使用。

一致的图像生成需在多张图像间忠实保留身份、风格和逻辑连贯性,这对故事叙述和角色设计至关重要。监督学习因缺乏大规模视觉一致性数据集及人类感知偏好建模复杂而受限。本文提出PaCo-RL,一种基于强化学习的框架,使模型能在无数据条件下学习复杂主观视觉标准。其核心为PaCo-Reward:一个通过自动化子图配对构建的大规模数据集训练的成对一致性评估器,采用生成式自回归打分机制,结合任务感知指令与思维链推理。另一组件PaCo-GRPO引入新颖的分辨率解耦优化策略,大幅降低强化学习开销,并采用日志截断多奖励聚合机制,确保奖励优化平衡稳定。在两个代表性子任务上的实验表明,PaCo-Reward显著提升与人类感知的一致性对齐,PaCo-GRPO实现当前最优一致性表现,同时具备更高训练效率与稳定性。结果验证了PaCo-RL作为可扩展、实用的一致图像生成方案的潜力。

原文摘要 · Abstract (English)

Consistent image generation requires faithfully preserving identities, styles, and logical coherence across multiple images, which is essential for applications such as storytelling and character design. Supervised training approaches struggle with this task due to the lack of large-scale datasets capturing visual consistency and the complexity of modeling human perceptual preferences. In this paper, we argue that reinforcement learning (RL) offers a promising alternative by enabling models to learn complex and subjective visual criteria in a data-free manner. To achieve this, we introduce PaCo-RL, a comprehensive framework that combines a specialized consistency reward model with an efficient RL algorithm. The first component, PaCo-Reward, is a pairwise consistency evaluator trained on a large-scale dataset constructed via automated sub-figure pairing. It evaluates consistency through a generative, autoregressive scoring mechanism enhanced by task-aware instructions and CoT reasons. The second component, PaCo-GRPO, leverages a novel resolution-decoupled optimization strategy to substantially reduce RL cost, alongside a log-tamed multi-reward aggregation mechanism that ensures balanced and stable reward optimization. Extensive experiments across the two representative subtasks show that PaCo-Reward significantly improves alignment with human perceptions of visual consistency, and PaCo-GRPO achieves state-of-the-art consistency performance with improved training efficiency and stability. Together, these results highlight the promise of PaCo-RL as a practical and scalable solution for consistent image generation. The project page is available at https://x-gengroup.github.io/HomePage_PaCo-RL/.

图像生成强化学习一致性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。