arXiv:2603.12247cs.CV2026-03被引 6

解决图像生成与编辑中奖励模型幻觉问题,提升真实性和指令遵循度。

Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation

  • 构建高质量评分数据集,分编辑与生成任务设计评估标准。
  • 训练出对齐人类判断的80亿参数奖励模型,显著减少幻觉。
  • 提出新奖励策略,适配不同任务,适合追求高精度生成的研究者。

强化学习(RL)在提升图像编辑和文本到图像(T2I)生成方面展现出巨大潜力。然而,当前作为批评者的奖励模型常因幻觉和噪声评分而误导优化过程。本文提出FIRM(Faithful Image Reward Modeling),一个全面框架,旨在构建鲁棒的奖励模型以提供准确可靠的指导。首先,设计定制化数据整理流程,分别基于执行一致性和指令遵循性构建编辑与生成评估标准。据此收集了FIRM-Edit-370K和FIRM-Gen-293K数据集,并训练出专用奖励模型(FIRM-Edit-8B和FIRM-Gen-8B)。其次,提出FIRM-Bench基准,专门用于评估编辑与生成中的批评者表现。实验显示,我们的模型在人类判断对齐上优于现有指标。为进一步融入RL流程,我们提出“基础+额外”奖励策略,平衡目标:编辑任务采用一致性调制执行(CME),生成任务采用质量调制对齐(QMA)。借助该框架,最终模型FIRM-Qwen-Edit与FIRM-SD3.5取得显著性能突破。实验证明FIRM有效抑制幻觉,确立了真实性和指令遵循的新标准。所有数据集、模型与代码已公开于https://firm-reward.github.io。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as a promising paradigm for enhancing image editing and text-to-image (T2I) generation. However, current reward models, which act as critics during RL, often suffer from hallucinations and assign noisy scores, inherently misguiding the optimization process. In this paper, we present FIRM (Faithful Image Reward Modeling), a comprehensive framework that develops robust reward models to provide accurate and reliable guidance for faithful image generation and editing. First, we design tailored data curation pipelines to construct high-quality scoring datasets. Specifically, we evaluate editing using both execution and consistency, while generation is primarily assessed via instruction following. Using these pipelines, we collect the FIRM-Edit-370K and FIRM-Gen-293K datasets, and train specialized reward models (FIRM-Edit-8B and FIRM-Gen-8B) that accurately reflect these criteria. Second, we introduce FIRM-Bench, a comprehensive benchmark specifically designed for editing and generation critics. Evaluations demonstrate that our models achieve superior alignment with human judgment compared to existing metrics. Furthermore, to seamlessly integrate these critics into the RL pipeline, we formulate a novel "Base-and-Bonus" reward strategy that balances competing objectives: Consistency-Modulated Execution (CME) for editing and Quality-Modulated Alignment (QMA) for generation. Empowered by this framework, our resulting models FIRM-Qwen-Edit and FIRM-SD3.5 achieve substantial performance breakthroughs. Comprehensive experiments demonstrate that FIRM mitigates hallucinations, establishing a new standard for fidelity and instruction adherence over existing general models. All of our datasets, models, and code have been publicly available at https://firm-reward.github.io.

图像生成强化学习奖励建模忠实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。