arXiv:2508.07818cs.CV2025-08被引 4

用视觉语言模型细粒度分析图像局部质量,提升感知准确性。

Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models

  • 用SAM分割图像为语义区域,再由多模态大模型分析每块质量
  • 提出区域感知注意力机制,融合局部特征生成全局质量判断
  • 不依赖特定主干网络,适配多种模型架构,提升泛化能力

无参考图像质量评估(NR-IQA)旨在模拟与人类主观感知一致的图像质量判断过程。然而,现有方法或仅关注全局表征,难以揭示语义显著区域;或对区域特征采用统一加权,削弱了对局部质量变化的敏感性。本文提出细粒度图像质量评估模型RSFIQA,通过整合区域级失真信息,感知多维度质量差异。首先利用分割一切模型(SAM)动态将输入图像划分为非重叠语义区域;对每个区域,训练强大的多模态大语言模型(MLLM)提取描述性内容并感知多维失真,实现对局部语义与质量退化的全面理解。为有效利用该信息,引入区域感知语义注意力(RSA)机制,通过聚合局部区域的细粒度表征生成全局注意力图。此外,RSFIQA具有主干无关性,可无缝集成至多种深度神经网络架构。大量实验表明,该方法在多个基准数据集上均展现出优异的鲁棒性与有效性。

原文摘要 · Abstract (English)

No-reference image quality assessment (NR-IQA) aims to simulate the process of perceiving image quality aligned with subjective human perception. However, existing NR-IQA methods either focus on global representations that leads to limited insights into the semantically salient regions or employ a uniform weighting for region features that weakens the sensitivity to local quality variations. In this paper, we propose a fine-grained image quality assessment model, named RSFIQA, which integrates region-level distortion information to perceive multi-dimensional quality discrepancies. To enhance regional quality awareness, we first utilize the Segment Anything Model (SAM) to dynamically partition the input image into non-overlapping semantic regions. For each region, we teach a powerful Multi-modal Large Language Model (MLLM) to extract descriptive content and perceive multi-dimensional distortions, enabling a comprehensive understanding of both local semantics and quality degradations. To effectively leverage this information, we introduce Region-Aware Semantic Attention (RSA) mechanism, which generates a global attention map by aggregating fine-grained representations from local regions. In addition, RSFIQA is backbone-agnostic and can be seamlessly integrated into various deep neural network architectures. Extensive experiments demonstrate the robustness and effectiveness of the proposed method, which achieves competitive quality prediction performance across multiple benchmark datasets.

图像质量评估多模态模型区域感知大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。