构建可扩展的人类对齐图像编辑评估基准,解决主观评价难题。
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
- 设计多维度自动评估管道,融合多种评分指标模拟人类感知。
- 覆盖广泛编辑任务的大规模数据集,支持可靠性能对比。
- 为先进模型提供基准结果,适合评估与改进文本引导图像编辑方法。
近期提出了多种文本引导图像编辑模型,但由于任务的主观性,尚无统一的标准评估方法,研究者普遍依赖人工用户研究。为此,我们提出一种新型的文本引导图像编辑人类对齐基准(HATIE)。该基准包含大规模、覆盖多种编辑任务的数据集,支持超越特定易评估案例的可靠评估。同时,HATIE提供完全自动化、全方位的评估流程,通过整合多种衡量编辑不同方面的评分,实现与人类感知的一致性。我们实证验证了HATIE评估结果在多个维度上确实与人类判断一致,并在多个前沿模型上提供了基准性能结果,深入揭示其表现差异。
原文摘要 · Abstract (English)
A variety of text-guided image editing models have been proposed recently. However, there is no widely-accepted standard evaluation method mainly due to the subjective nature of the task, letting researchers rely on manual user study. To address this, we introduce a novel Human-Aligned benchmark for Text-guided Image Editing (HATIE). Providing a large-scale benchmark set covering a wide range of editing tasks, it allows reliable evaluation, not limited to specific easy-to-evaluate cases. Also, HATIE provides a fully-automated and omnidirectional evaluation pipeline. Particularly, we combine multiple scores measuring various aspects of editing so as to align with human perception. We empirically verify that the evaluation of HATIE is indeed human-aligned in various aspects, and provide benchmark results on several state-of-the-art models to provide deeper insights on their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。