arXiv:2511.18055cs.CVcs.AI2025-11被引 1

提出新评估框架与模型,让图像编辑质量更贴合人眼判断。

IE-Critic-R1: Advancing the Explanatory Measurement of Text-Driven Image Editing for Human Perception Alignment

  • 基于可验证奖励的强化学习,提升评估可解释性
  • 在近4000张图像上实现与人类评分高度一致
  • 适合研究图像编辑评估或需感知对齐的场景

文本驱动图像编辑近年来发展迅速,但其质量评估仍面临挑战。与文本生成不同,编辑任务需同时依赖文本和源图,且结果随语义动态变化。现有方法多聚焦于文本-图像对齐,未能充分匹配人类感知。为此,本文构建了文本驱动图像编辑基准套件(IE-Bench),包含多样源图、多种编辑提示及对应结果,涵盖近4000个样本,并由15名受试者提供均值意见分(MOS)。同时提出IE-Critic-R1,借助可验证奖励的强化学习(RLVR),实现更全面且可解释的质量评估,显著优于以往指标,在主观感知对齐方面表现优异。相关数据与代码已公开。

原文摘要 · Abstract (English)

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image generation, text-driven image editing is characterized by simultaneously conditioning on both text and a source image. The edited images often retain an intrinsic connection to the original image, which dynamically change with the semantics of the text. However, previous methods tend to solely focus on text-image alignment or have not well aligned with human perception. In this work, we introduce the Text-driven Image Editing Benchmark suite (IE-Bench) to enhance the assessment of text-driven edited images. IE-Bench includes a database contains diverse source images, various editing prompts and the corresponding edited results from different editing methods, and nearly 4,000 samples with corresponding Mean Opinion Scores (MOS) provided by 15 human subjects. Furthermore, we introduce IE-Critic-R1, which, benefiting from Reinforcement Learning from Verifiable Rewards (RLVR), provides more comprehensive and explainable quality assessment for text-driven image editing that aligns with human perception. Extensive experiments demonstrate IE-Critic-R1's superior subjective-alignments on the text-driven image editing task compared with previous metrics. Related data and codes are available to the public.

图像编辑评估基准感知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。