arXiv:2603.14916cs.CVcs.MM2026-03被引 4

构建百万级图像编辑偏好数据集,实现更贴近人类审美的自动评估。

EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing

  • 基于多维度人类评分构建百万级图像编辑数据集
  • 提出多模态模型EditHF,对编辑结果进行人类对齐评估
  • 用该模型作为奖励信号优化编辑模型,提升真实感与指令遵循能力

近期文本引导的图像编辑(TIE)模型取得显著进展,但许多编辑结果仍存在伪影、意外修改和审美不佳等问题。尽管已有部分评测基准和方法,但缺乏可扩展的人类反馈评估模型,限制了图像编辑中人类偏好奖励模型的发展。为此,我们首次提出【EditHF-1M】——一个包含超过2900万组人类偏好对比和14.8万条人工平均评分的百万级图像编辑数据集,所有评分均从视觉质量、指令对齐度和属性保持性三个维度进行评估。基于此数据集,我们提出【EditHF】——一种基于多模态大语言模型的评估模型,可提供与人类偏好对齐的编辑反馈。进一步地,我们构建【EditHF-Reward】,利用EditHF作为奖励信号,通过强化学习优化文本引导图像编辑模型。大量实验表明,EditHF在人类偏好对齐方面表现优异,并在其他数据集上具备强泛化能力。我们还使用EditHF-Reward微调Qwen-Image-Edit,获得显著性能提升,证明其作为可扩展奖励模型的有效性。数据集与代码将开源于GitHub:https://github.com/IntMeGroup/EditHF。

原文摘要 · Abstract (English)

Recent text-guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthetic contents. Although some benchmarks and methods have been proposed for evaluating edited images, scalable evaluation models are still lacking, which limits the development of human feedback reward models for image editing. To address the challenges, we first introduce \textbf{EditHF-1M}, a million-scale image editing dataset with over 29M human preference pairs and 148K human mean opinion ratings, both evaluated from three dimensions, \textit{i.e.}, visual quality, instruction alignment, and attribute preservation. Based on EditHF-1M, we propose \textbf{EditHF}, a multimodal large language model (MLLM) based evaluation model, to provide human-aligned feedback from image editing. Finally, we introduce \textbf{EditHF-Reward}, which utilizes EditHF as the reward signal to optimize the text-guided image editing models through reinforcement learning. Extensive experiments show that EditHF achieves superior alignment with human preferences and demonstrates strong generalization on other datasets. Furthermore, we fine-tune the Qwen-Image-Edit using EditHF-Reward, achieving significant performance improvements, which demonstrates the ability of EditHF to serve as a reward model to scale-up the image editing. Both the dataset and code will be released in our GitHub repository: https://github.com/IntMeGroup/EditHF.

图像编辑人类偏好奖励建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。