用自编码器潜空间度量图像差异,更符合人眼感知。
Image Difference Quantification Using Autoencoder-Based Latent Representations
- 通过卷积自编码器提取图像潜变量,用余弦相似度量化差异。
- 狗猫图像对98.4%的相似度低于0.5,跨域数据集聚类清晰。
- 适合图像检索、质量评估等需感知一致性的场景。
传统图像相似性度量如均方误差(MSE)、峰值信噪比(PSNR)和结构相似性指数(SSIM)依赖像素级比较,常无法捕捉人眼感知的图像差异。相反,深度神经网络学习的潜空间表示包含高层语义信息,更贴近人类视觉感知。本文提出一种基于卷积自编码器的框架,利用潜空间中的余弦相似度量化图像差异。所学紧凑嵌入在光照、姿态和背景变化下仍能有效区分视觉差异显著的图像。在犬猫图像及多个跨域数据集上的广泛评估显示,潜空间中类别间聚类明显,类间可分性强,98.4%的犬猫图像对相似度低于0.5。进一步在TID2013数据集上的验证表明,潜空间距离与人类平均意见得分(MOS)正相关,对感知相关的图像失真敏感。该方法计算高效,语义基础扎实,可替代传统的像素级相似性度量,在基于内容的检索、感知质量评估和语义相似性分析中具有应用潜力。
原文摘要 · Abstract (English)
Traditional image similarity metrics such as Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), and the Structural Similarity Index Measure (SSIM) rely on pixel-level comparisons and often fail to capture perceptually meaningful differences between images. In contrast, latent representations learned by deep neural networks encode high-level semantic information that is more closely aligned with human visual perception. This paper proposes a convolutional autoencoder-based framework for quantifying image differences using cosine similarity in latent space. The learned compact embeddings enable robust differentiation between visually distinct images under variations in illumination, pose, and background. Extensive evaluation on dog-cat images and additional cross-domain datasets demonstrates clear class-wise clustering and strong inter-class separability in the latent space, with 98.4% of dog-cat image pairs exhibiting similarity scores below 0.5. Further validation using the TID2013 dataset shows that latent-space distance correlates positively with human Mean Opinion Scores (MOS), demonstrating sensitivity to perceptually relevant image distortions. The proposed approach provides a computationally efficient and semantically grounded alternative to conventional pixel-based similarity metrics, with potential applications in content-based retrieval, perceptual quality assessment, and semantic similarity analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。