arXiv:2505.15867cs.CVcs.LG2025-05ICML被引 2

不依赖标注数据,用图自编码器提升图像检索的语义准确性

SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval

  • 基于图自编码器实现无监督场景图检索,摆脱对标注数据依赖
  • 在多个指标上超越现有视觉、多模态及有监督方法,且运行更快
  • 首次引入图编辑距离作为评估标准,提升检索可靠性,适合语义检索研究者

尽管卷积和Transformer架构在图像到图像检索中占主导地位,但其易受颜色等低层视觉特征偏差影响。针对语义理解不足这一关键局限,本文提出SCENIR——一种强调语义内容而非表面特征的场景图检索框架。现有方法多依赖有监督图神经网络(GNN),需基于图像描述生成真实场景图对,但描述编码不一致导致监督信号不可靠。为此,我们设计基于图自编码器的无监督框架,完全消除对标签数据的需求。模型在各项指标和运行效率上均表现优异,超越现有视觉、多模态及有监督GNN方法。我们首次将图编辑距离(GED)作为确定性、鲁棒的场景图相似性评估标准,替代不一致的描述依赖。此外,通过自动构建场景图,将方法应用于未标注数据集,显著推动了反事实图像检索的性能边界。

原文摘要 · Abstract (English)

Despite the dominance of convolutional and transformer-based architectures in image-to-image retrieval, these models are prone to biases arising from low-level visual features, such as color. Recognizing the lack of semantic understanding as a key limitation, we propose a novel scene graph-based retrieval framework that emphasizes semantic content over superficial image characteristics. Prior approaches to scene graph retrieval predominantly rely on supervised Graph Neural Networks (GNNs), which require ground truth graph pairs driven from image captions. However, the inconsistency of caption-based supervision stemming from variable text encodings undermine retrieval reliability. To address these, we present SCENIR, a Graph Autoencoder-based unsupervised retrieval framework, which eliminates the dependence on labeled training data. Our model demonstrates superior performance across metrics and runtime efficiency, outperforming existing vision-based, multimodal, and supervised GNN approaches. We further advocate for Graph Edit Distance (GED) as a deterministic and robust ground truth measure for scene graph similarity, replacing the inconsistent caption-based alternatives for the first time in image-to-image retrieval evaluation. Finally, we validate the generalizability of our method by applying it to unannotated datasets via automated scene graph generation, while substantially contributing in advancing state-of-the-art in counterfactual image retrieval.

图像检索场景图无监督学习图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。