提出高效且抗干扰的图像质量评估模型,速度更快、更稳定。
BiRQA: Bidirectional Robust Quality Assessment for Images
- 双向多尺度金字塔结构融合四种快速特征,提升评估效率。
- 在五个基准上超越或持平当前最优,速度达前代3倍。
- 抗干扰能力显著增强,适合实时图像处理与生成应用。
全参考图像质量评估(FR IQA)对图像压缩、修复和生成建模至关重要,但现有神经度量方法仍存在速度慢、易受对抗扰动影响的问题。本文提出BiRQA,一种紧凑的全参考图像质量评估模型,通过双向多尺度金字塔并行处理四种快速互补特征。底层注意力模块通过不确定性感知门控将细粒度信息注入粗粒度层级,顶层交叉门控块则将语义上下文回传至高分辨率。为增强鲁棒性,引入锚定对抗训练策略,基于干净“锚点”样本与排序损失,理论上约束攻击下的逐点预测误差。在五个公开的FR IQA基准上,BiRQA性能优于或匹配现有最先进模型,且运行速度比前代SOTA快约3倍。在未见过的白盒攻击下,其KADID-10k数据集上的SROCC从0.30–0.57提升至0.60–0.84,展现出显著的鲁棒性提升。据我们所知,BiRQA是首个同时具备优异精度、实时吞吐能力和强对抗鲁棒性的全参考图像质量评估模型。
原文摘要 · Abstract (English)
Full-Reference image quality assessment (FR IQA) is important for image compression, restoration and generative modeling, yet current neural metrics remain slow and vulnerable to adversarial perturbations. We present BiRQA, a compact FR IQA metric model that processes four fast complementary features within a bidirectional multiscale pyramid. A bottom-up attention module injects fine-scale cues into coarse levels through an uncertainty-aware gate, while a top-down cross-gating block routes semantic context back to high resolution. To enhance robustness, we introduce Anchored Adversarial Training, a theoretically grounded strategy that uses clean "anchor" samples and a ranking loss to bound pointwise prediction error under attacks. On five public FR IQA benchmarks BiRQA outperforms or matches the previous state of the art (SOTA) while running ~3x faster than previous SOTA models. Under unseen white-box attacks it lifts SROCC from 0.30-0.57 to 0.60-0.84 on KADID-10k, demonstrating substantial robustness gains. To our knowledge, BiRQA is the only FR IQA model combining competitive accuracy with real-time throughput and strong adversarial resilience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。