多尺度注意力网络提升无参考图像质量评估,更准更高效。
MS-SCANet: A Multiscale Transformer-Based Architecture with Dual Attention for No-Reference Image Quality Assessment
- 多尺度双分支结构融合空间与通道注意力,捕获图像细节
- 在KonIQ-10k等数据集上相关性超现有方法,最高达0.947
- 适合图像质量分析、视觉感知优化等场景使用
我们提出多尺度空间通道注意力网络(MS-SCANet),一种基于Transformer的无参考图像质量评估(IQA)架构。该模型采用双分支结构,在多个尺度上处理图像,有效捕捉精细与粗略特征,优于传统单尺度方法。通过引入定制化空间和通道注意力机制,模型突出关键特征并降低计算复杂度。其核心是跨分支注意力机制,增强不同尺度间特征融合,解决以往方法的局限性。此外,我们设计了两种新一致性损失函数:跨分支一致性损失与自适应池化一致性损失,保持特征缩放过程中的空间完整性,优于传统的线性与双线性插值方法。在KonIQ-10k、LIVE、LIVE Challenge和CSIQ等数据集上的广泛实验表明,MS-SCANet持续超越现有最先进方法,与主观人类评分具有更强的相关性。
原文摘要 · Abstract (English)
We present the Multi-Scale Spatial Channel Attention Network (MS-SCANet), a transformer-based architecture designed for no-reference image quality assessment (IQA). MS-SCANet features a dual-branch structure that processes images at multiple scales, effectively capturing both fine and coarse details, an improvement over traditional single-scale methods. By integrating tailored spatial and channel attention mechanisms, our model emphasizes essential features while minimizing computational complexity. A key component of MS-SCANet is its cross-branch attention mechanism, which enhances the integration of features across different scales, addressing limitations in previous approaches. We also introduce two new consistency loss functions, Cross-Branch Consistency Loss and Adaptive Pooling Consistency Loss, which maintain spatial integrity during feature scaling, outperforming conventional linear and bilinear techniques. Extensive evaluations on datasets like KonIQ-10k, LIVE, LIVE Challenge, and CSIQ show that MS-SCANet consistently surpasses state-of-the-art methods, offering a robust framework with stronger correlations with subjective human scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。