提出评估图像生成中语义一致性与结构稳定性的新框架。
StableI2I: Spotting Unintended Changes in Image-to-Image Transition

- 通过动态分析输入输出图像的语义与空间一致性,无需参考图。
- 在多种图像编辑与修复任务中实现高精度、细粒度评估。
- 适合用于真实场景下模型性能诊断与对比评测。
在多数现实世界的图像到图像(I2I)任务中,现有评估方法主要关注指令遵循能力及生成图像的感知质量或美学效果,但严重忽略输出图像是否保持了输入图像的语义对应关系与空间结构。为解决这一问题,我们提出StableI2I——一个统一且动态的评估框架,可无须参考图像,在广泛I2I任务中显式衡量内容保真度与前后一致性,涵盖图像编辑与图像恢复等场景。此外,我们构建了StableI2I-Bench基准,系统评估多模态大模型(MLLMs)在保真度与一致性判断任务上的表现。大量实验表明,StableI2I能提供准确、细粒度且可解释的评估结果,与人类主观判断高度相关。该框架可作为实际可靠的工具,用于诊断内容一致性并评估真实I2I系统的模型性能。
原文摘要 · Abstract (English)
In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. However, they largely fail to assess whether the output image preserves the semantic correspondence and spatial structure of the input image. To address this limitation, we propose StableI2I, a unified and dynamic evaluation framework that explicitly measures content fidelity and pre--post consistency across a wide range of I2I tasks without requiring reference images, including image editing and image restoration. In addition, we construct StableI2I-Bench, a benchmark designed to systematically evaluate the accuracy of MLLMs on such fidelity and consistency assessment tasks. Extensive experimental results demonstrate that StableI2I provides accurate, fine-grained, and interpretable evaluations of content fidelity and consistency, with strong correlations to human subjective judgments. Our framework serves as a practical and reliable evaluation tool for diagnosing content consistency and benchmarking model performance in real-world I2I systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。