arXiv:2512.18407cs.CV2025-12

通过重要性感知图提升图像检索的语义准确性

Through the PRISm: Importance-Aware Scene Graphs for Image Retrieval

  • 基于重要性预测修剪无关对象,保留关键物体与关系
  • 在多个数据集上达到领先检索性能,接近人类感知水平
  • 适合需要精准语义检索的应用,如跨模态搜索

准确检索语义相似图像仍是计算机视觉中的基础挑战,传统方法常无法捕捉场景的关联与上下文细节。本文提出PRISm(基于剪枝的图像检索重要性预测语义图),通过两个新组件实现图像到图像检索的突破:首先,重要性预测模块识别并保留图像中最具语义重要性的物体和关系三元组,同时剪枝无关内容;其次,边感知图神经网络显式编码关系结构,并融合全局视觉特征,生成语义感知的图像嵌入。该框架通过显式建模物体及其交互的重要性,使检索结果更贴近人类感知,显著优于以往方法。其架构有效结合了关系推理与视觉表征,实现语义驱动的检索。在基准与真实世界数据集上的大量实验表明,该方法在顶级排名表现上持续领先;定性分析显示,PRISm能准确捕捉关键物体与交互,结果可解释且语义清晰。

原文摘要 · Abstract (English)

Accurately retrieving images that are semantically similar remains a fundamental challenge in computer vision, as traditional methods often fail to capture the relational and contextual nuances of a scene. We introduce PRISm (Pruning-based Image Retrieval via Importance Prediction on Semantic Graphs), a multimodal framework that advances image-to-image retrieval through two novel components. First, the Importance Prediction Module identifies and retains the most critical objects and relational triplets within an image while pruning irrelevant elements. Second, the Edge-Aware Graph Neural Network explicitly encodes relational structure and integrates global visual features to produce semantically informed image embeddings. PRISm achieves image retrieval that closely aligns with human perception by explicitly modeling the semantic importance of objects and their interactions, capabilities largely absent in prior approaches. Its architecture effectively combines relational reasoning with visual representation, enabling semantically grounded retrieval. Extensive experiments on benchmark and real-world datasets demonstrate consistently superior top-ranked performance, while qualitative analyses show that PRISm accurately captures key objects and interactions, producing interpretable and semantically meaningful results.

图像检索语义图关系推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。