用流形几何优化原型匹配,让细粒度识别更可解释且更准确。
Geodesic Prototype Matching via Diffusion Maps for Interpretable Fine-Grained Recognition
- 基于扩散映射构建类内流形结构,用内在几何定义相似性。
- 在两个数据集上超越传统欧氏距离原型网络,显著提升识别精度。
- 适合需要可解释性的细粒度分类场景,如医学图像分析。
深层视觉特征中普遍存在非线性流形,欧氏距离会扭曲真实相似性,这对依赖原型的可解释细粒度识别尤为不利,因细微语义差异至关重要。本文提出一种新范式,将相似性建立在深度特征的内在几何基础上。具体而言,将每类的潜在流形结构提炼为扩散空间,并设计可微分的Nyström插值,使该几何结构对未见样本和可学习原型均可用。为保持效率,采用周期性更新的紧凑类级地标集,确保嵌入与不断演化的主干网络同步,支持大规模快速推理。在两个基准数据集上的全面实验表明,所提出的GeoProto能生成聚焦于语义对应部位的原型,显著优于欧氏原型网络。
原文摘要 · Abstract (English)
Nonlinear manifolds are pervasive in deep visual features, where Euclidean distances can misrepresent true similarity. This mismatch is particularly detrimental to prototype-based interpretable fine-grained recognition, where even subtle semantic distinctions are crucial. To mitigate this issue, this work presents a novel paradigm for prototype-based recognition by grounding similarity in the intrinsic geometry of deep features. Concretely, we distill the latent manifold structure of each class into a diffusion space and, critically, devise a differentiable Nyström interpolation to make this geometry accessible to both unseen samples and learnable prototypes. To maintain efficiency, we employ compact per-class landmark sets with periodic updates. This strategy keeps the embedding synchronized with the evolving backbone, enabling fast inference at scale. Comprehensive experiments on two benchmark datasets demonstrate that our GeoProto yields prototypes focusing on semantically corresponding parts, significantly outperforming Euclidean prototype networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。