arXiv:2510.18377cs.CV2025-10

用图文语义对齐提升图像复杂度评估,更贴近人眼感知。

Cross-Modal Scene Semantic Alignment for Image Complexity Assessment

  • 通过图文配对学习对齐图像与文本的场景语义信息。
  • 在多个数据集上超越现有方法,显著提升复杂度预测准确率。
  • 适合关注感知理解与跨模态融合的研究者和应用开发人员。

图像复杂度评估(ICA)因人类感知的主观性和真实世界图像的语义多样性而极具挑战。现有方法主要依赖单一视觉模态的手工特征或浅层卷积神经网络特征,难以充分捕捉与图像复杂度相关的感知表征。近期研究表明,跨模态场景语义信息在多种计算机视觉任务中起关键作用,尤其在感知理解方面。然而,其在ICA中的应用尚未被探索。为此,本文提出一种名为跨模态场景语义对齐(CM-SSA)的新方法,从跨模态视角利用场景语义对齐增强ICA性能,使复杂度预测更符合主观人类感知。具体而言,CM-SSA包含复杂度回归分支和场景语义对齐分支:前者在后者引导下估计图像复杂度等级,后者通过成对学习将图像与蕴含丰富场景语义的文本提示对齐。大量实验表明,所提方法在多个ICA数据集上显著优于当前最优方法。代码已开源:https://github.com/XQ2K/First-Cross-Model-ICA。

原文摘要 · Abstract (English)

Image complexity assessment (ICA) is a challenging task in perceptual evaluation due to the subjective nature of human perception and the inherent semantic diversity in real-world images. Existing ICA methods predominantly rely on hand-crafted or shallow convolutional neural network-based features of a single visual modality, which are insufficient to fully capture the perceived representations closely related to image complexity. Recently, cross-modal scene semantic information has been shown to play a crucial role in various computer vision tasks, particularly those involving perceptual understanding. However, the exploration of cross-modal scene semantic information in the context of ICA remains unaddressed. Therefore, in this paper, we propose a novel ICA method called Cross-Modal Scene Semantic Alignment (CM-SSA), which leverages scene semantic alignment from a cross-modal perspective to enhance ICA performance, enabling complexity predictions to be more consistent with subjective human perception. Specifically, the proposed CM-SSA consists of a complexity regression branch and a scene semantic alignment branch. The complexity regression branch estimates image complexity levels under the guidance of the scene semantic alignment branch, while the scene semantic alignment branch is used to align images with corresponding text prompts that convey rich scene semantic information by pair-wise learning. Extensive experiments on several ICA datasets demonstrate that the proposed CM-SSA significantly outperforms state-of-the-art approaches. Codes are available at https://github.com/XQ2K/First-Cross-Model-ICA.

图像评估跨模态语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。