arXiv:2507.17182cs.CV2025-07

多层级特征融合提升AIGC图像质量评估精度

Hierarchical Fusion and Joint Aggregation: A Multi-Level Feature Representation Method for AIGC Image Quality Assessment

  • 分三阶段构建多层级视觉表征,融合全局与局部特征
  • 在多个数据集上同时实现感知质量和图文一致性评估领先
  • 适合需要精准评估AI生成图像质量的研究者和开发者

AI生成内容(AIGC)的质量评估面临从低层视觉感知到高层语义理解的多重挑战。现有方法通常依赖单层次视觉特征,难以捕捉AIGC图像中的复杂失真。为此,提出一种多层级视觉表征范式,包含三个阶段:多层级特征提取、层级融合与联合聚合。基于该范式,设计两种网络:用于感知质量评估的多层级全局-局部融合网络(MGLF-Net),通过双卷积神经网络与视觉变换器骨干提取互补的局部与全局特征;用于文本到图像对应关系评估的多层级提示嵌入融合网络(MPEF-Net),在每一特征层级将提示语义嵌入视觉特征融合过程。融合后的多层级特征被联合聚合以进行最终评估。在多个基准数据集上的实验表明,该方法在两项任务上均表现卓越,验证了所提多层级视觉评估范式的有效性。

原文摘要 · Abstract (English)

The quality assessment of AI-generated content (AIGC) faces multi-dimensional challenges, that span from low-level visual perception to high-level semantic understanding. Existing methods generally rely on single-level visual features, limiting their ability to capture complex distortions in AIGC images. To address this limitation, a multi-level visual representation paradigm is proposed with three stages, namely multi-level feature extraction, hierarchical fusion, and joint aggregation. Based on this paradigm, two networks are developed. Specifically, the Multi-Level Global-Local Fusion Network (MGLF-Net) is designed for the perceptual quality assessment, extracting complementary local and global features via dual CNN and Transformer visual backbones. The Multi-Level Prompt-Embedded Fusion Network (MPEF-Net) targets Text-to-Image correspondence by embedding prompt semantics into the visual feature fusion process at each feature level. The fused multi-level features are then aggregated for final evaluation. Experiments on benchmarks demonstrate outstanding performance on both tasks, validating the effectiveness of the proposed multi-level visual assessment paradigm.

AIGC评估多层级特征图像质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。