提出PixLens框架,用检测+SAM实现扩散模型图像编辑的解耦评估
PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM
- 结合目标检测与SAM,自动定位编辑区域并评估效果
- 首次同时量化编辑质量与潜在表示解耦程度
- 适合研究生成模型评估与可控编辑的开发者
评估基于扩散的图像编辑模型是生成式AI领域的重要任务。必须评估其在保持图像内容和真实感的前提下执行多种编辑任务的能力。尽管生成模型的发展带来了前所未有的图像编辑可能性,但对这些模型进行全面评估仍具挑战性且尚未标准化。主要困难在于缺乏后编辑参考图像,导致评价依赖于CLIP等预训练模型或人工判断。本文提出的PixLens基准,能全面评估编辑质量与潜在表示的解耦程度,推动该领域方法论的发展与优化。
原文摘要 · Abstract (English)
Evaluating diffusion-based image-editing models is a crucial task in the field of Generative AI. Specifically, it is imperative to assess their capacity to execute diverse editing tasks while preserving the image content and realism. While recent developments in generative models have opened up previously unheard-of possibilities for image editing, conducting a thorough evaluation of these models remains a challenging and open task. The absence of a standardized evaluation benchmark, primarily due to the inherent need for a post-edit reference image for evaluation, further complicates this issue. Currently, evaluations often rely on established models such as CLIP or require human intervention for a comprehensive understanding of the performance of these image editing models. Our benchmark, PixLens, provides a comprehensive evaluation of both edit quality and latent representation disentanglement, contributing to the advancement and refinement of existing methodologies in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。