arXiv:2608.06142cs.CV2026-08

提出两种艺术构图分析方法,兼顾效果与可解释性。

Learning visual representations for compositional analysis of artworks and photographs

论文配图:Learning visual representations for compositional analysis of artworks and photographs
图 1 · 摘自论文原文
  • 用对象中心模型分解区域,图注意力网络捕捉元素空间关系。
  • 在构图评分、检索和显著性检测任务上表现优异,数据充足时大模型更优。
  • 适合关注可解释性的视觉理解研究者使用。

构图是艺术作品中意义、情感和美学质量传达的核心,但仍是视觉理解中最不规范的维度之一。现有研究指出,学习有意义的构图表征存在持续差距,归因于语义偏差,认为人类启发的方法可能是关键。本文对比两种构图分析范式:一种基于感知分组的人类启发方法,另一种由大规模构图数据集驱动的微调基础模型。前者采用对象中心模型进行区域级分解,并用图注意力网络捕捉元素间的空间关系;后者则利用自监督大模型进行微调。两者在构图评分/类别预测、构图图像检索和视觉显著性检测任务上进行评估。结果显示,在编码器冻结条件下,人类启发方法表现竞争性且可解释性强;当数据充足时,大模型显著超越,但牺牲了可解释性和跨域泛化能力。代码与预训练模型已开源。

原文摘要 · Abstract (English)

Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization. Code and pre-trained models are available on GitHub.

构图分析可解释性图神经网络视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。