arXiv:2604.19406cs.CVcs.AI2026-04中稿 · CVPR被引 4

用真人偏好数据提升图像编辑模型质量,让生成结果更符合人类审美。

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

论文配图:HP-Edit: A Human-Preference Post-Training Framework for Image Editing
图 1 · 摘自论文原文
  • 基于少量真人评分数据和视觉大模型,自动构建偏好评估器。
  • 在8类常见编辑任务中,显著提升模型与人类偏好的一致性。
  • 适合需要高质量图像编辑的AI研究者与产品开发者。

常见图像编辑任务通常采用强大的生成式扩散模型作为主流方案。尽管强化学习(如Diffusion-DPO和Flow-GRPO)进一步提升了生成质量,但由于缺乏可扩展的人类偏好数据集和针对多样化编辑需求的框架,将人类反馈强化学习(RLHF)应用于扩散模型编辑仍处于探索阶段。为此,我们提出HP-Edit——一种面向人类偏好对齐的后训练框架,并引入RealPref-50K,一个涵盖八类常见任务、平衡常见物体编辑的真实世界数据集。HP-Edit利用少量人类偏好评分数据和预训练视觉大语言模型(VLM),构建自动化的、与人类偏好对齐的评估器HP-Scorer。该评估器既可用于高效构建可扩展的偏好数据集,也可作为后训练编辑模型的奖励函数。我们还提出了RealPref-Bench,用于评估真实场景下的编辑性能。大量实验表明,该方法显著提升了Qwen-Image-Edit-2509等模型的表现,使其输出更贴近人类偏好。

原文摘要 · Abstract (English)

Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Flow-GRPO have further improved generation quality, efficiently applying Reinforcement Learning from Human Feedback (RLHF) to diffusion-based editing remains largely unexplored, due to a lack of scalable human-preference datasets and frameworks tailored to diverse editing needs. To fill this gap, we propose HP-Edit, a post-training framework for Human Preference-aligned Editing, and introduce RealPref-50K, a real-world dataset across eight common tasks and balancing common object editing. Specifically, HP-Edit leverages a small amount of human-preference scoring data and a pretrained visual large language model (VLM) to develop HP-Scorer--an automatic, human preference-aligned evaluator. We then use HP-Scorer both to efficiently build a scalable preference dataset and to serve as the reward function for post-training the editing model. We also introduce RealPref-Bench, a benchmark for evaluating real-world editing performance. Extensive experiments demonstrate that our approach significantly enhances models such as Qwen-Image-Edit-2509, aligning their outputs more closely with human preference.

图像编辑人类偏好扩散模型RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。