无需人工标注,通过结构化关联自监督学习图像质量评估。
SHAMISA: SHAped Modeling of Implicit Structural Associations for Self-supervised No-Reference Image Quality Assessment
- 利用合成元数据与特征结构推断软性、可调控的隐式关联关系。
- 在多个数据集上达到领先性能,且跨数据集泛化能力更强。
- 适合无标注图像质量评估研究者与工业部署场景使用。
无参考图像质量评估(NR-IQA)旨在不依赖原始参考图像的情况下估计感知质量。现有方法面临核心瓶颈:需大量昂贵的人类感知标注。本文提出SHAMISA,一种非对比性自监督框架,通过显式结构化关系监督从无标签失真图像中学习。不同于以往施加刚性二值相似性约束的方法,SHAMISA引入隐式结构关联——一种受失真感知与内容敏感性影响的软性、可控关系,由合成元数据和内在特征结构推导得出。关键创新在于组合式失真引擎,能从连续参数空间生成不可数数量的退化类型,每组仅单一退化因子变化,实现嵌入空间中表示相似性的精细控制:具有相同退化模式的图像被拉近,严重程度变化则产生结构化、可预测的偏移。通过双源关系图融合已知退化特征与涌现的结构亲和性,指导全程训练。采用卷积编码器在该监督下训练后冻结,质量预测由其特征上的线性回归器完成。大量实验表明,SHAMISA在合成、真实及跨数据集的NR-IQA基准测试中表现优异,具备更强跨数据集泛化能力与鲁棒性,且无需人类质量标注或对比损失。
原文摘要 · Abstract (English)
No-Reference Image Quality Assessment (NR-IQA) aims to estimate perceptual quality without access to a reference image of pristine quality. Learning an NR-IQA model faces a fundamental bottleneck: its need for a large number of costly human perceptual labels. We propose SHAMISA, a non-contrastive self-supervised framework that learns from unlabeled distorted images by leveraging explicitly structured relational supervision. Unlike prior methods that impose rigid, binary similarity constraints, SHAMISA introduces implicit structural associations, defined as soft, controllable relations that are both distortion-aware and content-sensitive, inferred from synthetic metadata and intrinsic feature structure. A key innovation is our compositional distortion engine, which generates an uncountable family of degradations from continuous parameter spaces, grouped so that only one distortion factor varies at a time. This enables fine-grained control over representational similarity during training: images with shared distortion patterns are pulled together in the embedding space, while severity variations produce structured, predictable shifts. We integrate these insights via dual-source relation graphs that encode both known degradation profiles and emergent structural affinities to guide the learning process throughout training. A convolutional encoder is trained under this supervision and then frozen for inference, with quality prediction performed by a linear regressor on its features. Extensive experiments on synthetic, authentic, and cross-dataset NR-IQA benchmarks demonstrate that SHAMISA achieves strong overall performance with improved cross-dataset generalization and robustness, all without human quality annotations or contrastive losses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。