让图像相似度评估更懂人眼偏好,提升文本生成图像的迭代体验。
CLPIPS: A Personalized Metric for AI-Generated Image Similarity

- 基于LPIPS改进,通过少量人类反馈微调权重,让评估更贴近人眼判断。
- 在人类评分数据集上,与原始LPIPS相比相关性提升显著。
- 适合需要人机协同优化的文生图工作流,尤其关注主观感知一致性。
迭代提示优化是使用文本到图像生成模型复现目标图像的核心。以往研究将图像相似度度量(ISMs)作为对人类用户的额外反馈。现有度量如LPIPS和CLIP虽提供客观图像相似性判断,但在上下文特定或用户驱动任务中常与人类判断不符。本文提出定制化学习感知图像块相似度(CLPIPS),是对LPIPS的定制扩展,可直接适应人类判断。我们探索轻量级、人类辅助微调是否能显著提升感知一致性,使相似度度量成为人机协同文生图流程中的可适配组件。在包含参与者反复生成目标图像并按感知相似性排序的人类实验数据集上进行评估。采用边缘排序损失对人类排名图像对进行微调,仅调整LPIPS层组合权重,并通过斯皮尔曼秩相关系数和组内相关系数评估对齐效果。结果表明,CLPIPS相较于基线LPIPS在与人类判断的一致性上表现更优。本工作强调的是提升度量预测与人类排名之间的一致性,而非绝对性能优化,证明即使有限的人类特异性微调也能显著增强人机协同文生图流程中的感知对齐。
原文摘要 · Abstract (English)
Iterative prompt refinement is central to reproducing target images with text to image generative models. Previous studies have incorporated image similarity metrics (ISMs) as additional feedback to human users. Existing ISMs such as LPIPS and CLIP provide objective measures of image likeness but often fail to align with human judgments, particularly in context specific or user driven tasks. In this paper, we introduce Customized Learned Perceptual Image Patch Similarity (CLPIPS), a customized extension of LPIPS that adapts a metric's notion of similarity directly to human judgments. We aim to explore whether lightweight, human augmented fine tuning can meaningfully improve perceptual alignment, positioning similarity metrics as adaptive components for human in the loop workflows with text to image tools. We evaluate CLPIPS on a human subject dataset in which participants iteratively regenerate target images and rank generated outputs by perceived similarity. Using margin ranking loss on human ranked image pairs, we fine tune only the LPIPS layer combination weights and assess alignment via Spearman rank correlation and Intraclass Correlation Coefficient. Our results show that CLPIPS achieves stronger correlation and agreement with human judgments than baseline LPIPS. Rather than optimizing absolute metric performance, our work emphasizes improving alignment consistency between metric predictions and human ranks, demonstrating that even limited human specific fine tuning can meaningfully enhance perceptual alignment in human in the loop text to image workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。