arXiv:2512.23169cs.CV2025-12ACL被引 3

用强化学习提升图文对齐评估的精细度和可解释性

REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation

  • 通过结构化推理流程让模型定位语义元素并生成判断
  • 在四个基准上超越主流模型,推理效率更高
  • 适合需要精准评估生成图像质量的研究与应用

评估文本提示与生成图像之间的对齐程度对保障文生图模型的可靠性与可用性至关重要。然而,现有方法多依赖粗粒度指标或静态问答流程,缺乏细粒度可解释性,难以反映人类偏好。为此,我们提出REVEALER,一种基于强化引导视觉推理的元素级对齐评估统一框架。采用“定位-推理-结论”的结构化范式,使多模态大语言模型能够显式定位语义元素并生成可解释的对齐判断。通过结合结构格式、定位准确率和对齐保真度的复合奖励函数,利用分组相对策略优化(GRPO)进行模型优化。在EvalMuse-40K、RichHF、MHaluBench和GenAI-Bench四个基准上的大量实验表明,REVEALER达到当前最优性能,显著优于强大多数专有模型与监督基线,并展现出相比现有迭代式视觉推理方法更优的推理效率。

原文摘要 · Abstract (English)

Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of text-to-image (T2I) models. However, most existing evaluation methods rely on coarse-grained metrics or static QA pipelines, which lack fine-grained interpretability and struggle to reflect human preferences. To address this, we propose REVEALER, a unified framework for element-level alignment evaluation based on reinforcement-guided visual reasoning. Adopting a structured "grounding-reasoning-conclusion" paradigm, our method enables Multimodal Large Language Models (MLLMs) to explicitly localize semantic elements and derive interpretable alignment judgments. We optimize the model via Group Relative Policy Optimization(GRPO) using a composite reward function that incorporates structural format, grounding accuracy, and alignment fidelity. Extensive experiments across four benchmarks-EvalMuse-40K, RichHF, MHaluBench, and GenAI-Bench-demonstrate that REVEALER achieves state-of-the-art performance. Our approach consistently outperforms both strong proprietary models and supervised baselines while demonstrating superior inference efficiency compared to existing iterative visual reasoning methods.

图文对齐视觉推理强化学习评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。