提出可解释的低维表示,提升图像检索与图神经网络分类效果
Context-Aware Interpretable Representations for Retrieval and Graph Convolutional Network Classification

- 基于流形学习与排序嵌入,融合上下文信息生成稀疏可解释表示
- 在多个数据集上实现图像检索准确率提升,半监督分类精度超过基线模型
- 适合需要模型可解释性与高效特征表示的研究者使用
近年来视觉信息建模和表示学习取得显著进展,主要得益于卷积神经网络、基于Transformer的模型及基础模型。然而,相似性评估的本质与模型透明性等关键问题仍被忽视。主要挑战包括:几何差距——传统成对度量无法捕捉数据集流形的内在结构;可解释性差距——表示与人类认知不一致。如何在保持低维性和下游任务有效性的同时提供可解释性,仍是开放问题。本文提出一种新颖的无监督框架,结合流形学习与基于排序的可解释图嵌入。该方法首先通过流形分析刻画数据集的上下文信息,再生成稀疏且自解释的嵌入表示。所提方法具有灵活架构,支持多种流形学习与表示学习策略。在多样化数据集与特征上的大量实验表明,其上下文感知表示不仅具备内在可解释性与降维能力,还在图像检索与图卷积网络(GCN)半监督分类任务中保持或提升了性能。
原文摘要 · Abstract (English)
The advances in visual information modeling and representation during the last decades are remarkable, mainly supported by Convolutional Neural Networks, Transformer-based, and Foundation Models. Despite this progress, critical challenges regarding the nature of similarity assessment and model transparency have been neglected. A primary concern is the Geometric Gap, where traditional pairwise measures fail to capture the intrinsic geometry of the dataset manifold. Furthermore, the Interpretability Gap persists, as representations often lack alignment with human cognition. Therefore, how to provide interpretability to representations while maintaining low dimensionality and high effectiveness in downstream tasks remains an open challenge. In this paper, we propose a novel unsupervised framework that integrates Manifold Learning strategies with Rank-based Interpretable Graph Embeddings. Our approach effectively bridges these gaps by first characterizing the contextual information of the dataset through manifold analysis and subsequently generating sparse, self-explainable embeddings. The proposed approach employs a flexible formulation, allowing different Manifold Learning and Representation Learning strategies. Extensive experimental evaluation across diverse datasets and features demonstrates that our Context-Aware representations not only provide intrinsic interpretability and dimensionality reduction but also maintain or enhance effectiveness in downstream tasks, specifically in image retrieval and semi-supervised classification using Graph Convolutional Networks (GCNs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。