多尺度特征提升图像相似性评估,效果显著且开销极小
MSDS: Deep Structural Similarity with Multiscale Representation

- 在不同分辨率层级独立计算相似性,再加权融合
- 多个数据集上相比单尺度模型性能显著提升
- 适合需要高精度图像质量评估的研究与应用
基于深度特征的感知相似性模型在图像质量评估中表现出与人类视觉感知高度一致。然而,现有方法大多仅在单一空间尺度上运行,隐含假设固定分辨率下的结构相似性已足够。空间尺度在深度特征相似性建模中的作用尚未充分理解。本文通过最小化多尺度扩展的DeepSSIM,提出多尺度结构相似性(MSDS)框架。该方法通过在金字塔层级上独立计算DeepSSIM,并使用可学习的轻量级全局权重融合得分,实现深度特征表示与跨尺度整合的解耦。在多个基准数据集上的实验表明,相较于单尺度基线,性能持续且统计显著提升,同时引入的额外复杂度可忽略不计。结果实证确认了空间尺度是深度感知相似性中不可忽视的因素,本研究以最小测试平台将其孤立验证。
原文摘要 · Abstract (English)
Deep-feature-based perceptual similarity models have demonstrated strong alignment with human visual perception in Image Quality Assessment (IQA). However, most existing approaches operate at a single spatial scale, implicitly assuming that structural similarity at a fixed resolution is sufficient. The role of spatial scale in deep-feature similarity modeling thus remains insufficiently understood. In this letter, we isolate spatial scale as an independent factor using a minimal multiscale extension of DeepSSIM, referred to as Deep Structural Similarity with Multiscale Representation (MSDS). The proposed framework decouples deep feature representation from cross-scale integration by computing DeepSSIM independently across pyramid levels and fusing the resulting scores with a lightweight set of learnable global weights. Experiments on multiple benchmark datasets demonstrate consistent and statistically significant improvements over the single-scale baseline, while introducing negligible additional complexity. The results empirically confirm spatial scale as a non-negligible factor in deep perceptual similarity, isolated here via a minimal testbed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。