arXiv:2510.12784cs.CVcs.CL2025-10中稿 · ECCV被引 17

让视觉模型自己评价并改进生成结果,无需人工标注。

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

  • 用理解模块当内部评分器,反哺生成模块。
  • 在多个基准上生成质量提升,最高达88.37分。
  • 适合想提升图像生成能力的统一多模态模型使用者。

统一多模态模型(UMMs)在视觉-语言生成与理解任务中取得显著进展,但其强大的视觉理解能力往往无法有效迁移到视觉生成:模型可能正确判断提示与图像的一致性,却无法从相同提示生成忠实图像。这引发了一个关键问题:能否让模型利用自身的理解模块来奖励其生成模块?我们提出SRUM,一种可直接应用于各类现有UMMs的自奖励后训练框架。SRUM构建了反馈环路,使模型的理解模块充当内部评估器,提供修正信号以提升生成效果,无需额外人工标注数据或外部奖励模型。为实现全面反馈,SRUM采用全局-局部双奖励机制:全局奖励确保整体语义与布局一致性,局部奖励细化物体级别的细节保真度。实验显示,SRUM在T2I-CompBench上性能从82.18提升至88.37,在T2I-ReasonBench上从43.82提升至46.75。本工作确立了一种新范式,使UMM的理解模块能够通过自奖励引导并增强自身生成能力。

原文摘要 · Abstract (English)

Recently, remarkable progress has been made in Unified Multimodal Models (UMMs), which integrate vision-language generation and understanding capabilities within a single framework. However, a model's strong visual understanding often fails to transfer to visual generation: it may correctly judge prompt-image alignment while failing to generate a faithful image from the same prompt. This raises a compelling question: Can a model improve itself by using its understanding module to reward its generation module? We introduce SRUM, a self-rewarding post-training framework directly applicable to existing UMMs of various designs. SRUM creates a feedback loop where the model's own understanding module acts as an internal ``evaluator'', providing corrective signals to improve generation without additional human-labeled data or external reward models. To provide comprehensive feedback, SRUM uses a global-local dual reward system: a \textbf{global reward} ensures overall visual semantics and layout, while a \textbf{local reward} refines fine-grained, object-level fidelity. SRUM shows strong generalization, boosting performance on T2I-CompBench from 82.18 to \textbf{88.37} and on T2I-ReasonBench from 43.82 to \textbf{46.75}. Overall, our work establishes a powerful paradigm for enabling a UMM's understanding module to guide and enhance its own generation via self-rewarding.

多模态自奖励图像生成统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。