arXiv:2501.15890cs.CVcs.AI2025-01被引 2

提出新方法解析视觉复杂度中的结构、色彩与意外感因素

Complexity in Complexity: Understanding Visual Complexity Through Structure, Color, and Surprise

  • 用多尺度梯度、色彩多样性和大模型惊喜度评估复杂度
  • 在新数据集上显著提升预测准确率,突破原有简单模型局限
  • 适合关注感知机制与可解释性的视觉认知研究者

理解人类对视觉复杂度的感知是视觉认知的关键课题。以往建模方法常依赖复杂算法或深度网络,虽在特定数据集表现优异,却牺牲了可解释性。近期(Shen等,2024)提出一种基于分割的可解释模型,能跨数据集准确预测复杂度,暗示复杂度或可简单解释。本文探究该模型未能捕捉结构、色彩与意外感贡献的原因,提出多尺度Sobel梯度(MSG)衡量空间强度变化,多尺度独特色彩(MUC)量化多尺度色度丰富度,并利用大语言模型生成惊喜度评分。我们在现有基准和新构建的数据集(Surprising Visual Genome)上测试,该数据集包含来自Visual Genome的令人意外图像。实验表明,准确建模复杂度远非简单,需引入额外感知与语义因素以应对数据集偏差。所提模型在保持可解释性的同时提升预测性能,深化了对视觉复杂度感知机制的理解。代码、分析与数据已开源。

原文摘要 · Abstract (English)

Understanding how humans perceive visual complexity is a key area of study in visual cognition. Previous approaches to modeling visual complexity assessments have often resulted in intricate, difficult-to-interpret algorithms that employ numerous features or sophisticated deep learning architectures. While these complex models achieve high performance on specific datasets, they often sacrifice interpretability, making it challenging to understand the factors driving human perception of complexity. Recently (Shen, et al. 2024) proposed an interpretable segmentation-based model that accurately predicted complexity across various datasets, supporting the idea that complexity can be explained simply. In this work, we investigate the failure of their model to capture structural, color and surprisal contributions to complexity. To this end, we propose Multi-Scale Sobel Gradient (MSG) which measures spatial intensity variations, Multi-Scale Unique Color (MUC) which quantifies colorfulness across multiple scales, and surprise scores generated using a Large Language Model. We test our features on existing benchmarks and a novel dataset (Surprising Visual Genome) containing surprising images from Visual Genome. Our experiments demonstrate that modeling complexity accurately is not as simple as previously thought, requiring additional perceptual and semantic factors to address dataset biases. Our model improves predictive performance while maintaining interpretability, offering deeper insights into how visual complexity is perceived and assessed. Our code, analysis and data are available at https://github.com/Complexity-Project/Complexity-in-Complexity.

视觉认知复杂度建模可解释性大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。