arXiv:2607.16321cs.CVcs.IR2026-07中稿 · ECCV被引 1

用层化理论构建多关系艺术表示,让视觉与文本对齐更贴近艺术史分析。

Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations

论文配图:Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
图 1 · 摘自论文原文
  • 基于层化理论,为每幅画生成多种上下文相关嵌入
  • 在三个新数据集上显著优于基线模型,提升跨模态检索效果
  • 无需外部数据即可推理,适合艺术史与多视角理解任务

理解一幅画从来不是单一行为。艺术史学家可能从风格、图像志或历史语境等维度分析同一作品,这些维度不可互换,且在视觉与文本间存在不同语义关系。像CLIP这样的视觉语言模型通过单一共享嵌入空间学习对齐,将这种丰富性压缩为同质化表示,从而丢失了艺术史推理所依赖的多关系结构。我们提出CANVAS(Contrastive Art-aware Network for Vision-Language Alignment with Sheaves),一种受层化理论启发的框架,用于学习关系感知的多模态表示。每幅作品被映射到多个基于关系类型(即上下文)的嵌入,并采用新颖对比损失在训练中编码上下文信息,推理阶段无需依赖外部数据。我们在三个新引入的艺术多关系理解基准上进行评估:源自WikiArt和Wikipedia的WikiArt+,来自Bibliotheca Hertziana收藏的HertzianaDP,以及经精炼的SemArt+。在跨模态检索与艺术理解任务中,CANVAS表现优于基线,支持多关系对齐不仅是理论必要,更是实际关键。

原文摘要 · Abstract (English)

Understanding a painting is never a single act. Art historians may analyze the same work through concepts of style, iconography, or historical context, dimensions that are not interchangeable, and each carries distinct semantic relationships between the visual and the textual. Vision-Language Models (VLMs) like CLIP, which learn a single shared embedding space, collapse this richness into a single homogeneous alignment, thereby losing the multi-relational structure that defines art-historical reasoning. We introduce CANVAS (Contrastive Art-aware Network for Vision-Language Alignment with Sheaves), a framework for learning relation-aware multimodal representations inspired by sheaf theory. Each artwork is projected into multiple embeddings conditioned on the type of relation (i.e., the context), and a novel contrastive loss encodes contextual information during training, with no dependency on external data at inference. We evaluate on three newly introduced benchmarks of artworks for multi-relational art understanding: WikiArt+, derived from WikiArt and Wikipedia, HertzianaDP, from the Bibliotheca Hertziana collection, and SemArt+, refined from the SemArt dataset. In multimodal retrieval and art understanding, CANVAS outperforms the baselines, supporting the view that multi-relational alignment is not just theoretically motivated but also practically essential.

艺术理解多关系层化理论对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。